Archive / 2026-09-08
September 8, 2026
News
View all news →mistral.ai10 days ago841 ptsView detailsJoin discussion
LibreOffice breaks download records after declaring it has no AI features
manualdousuario.net9 days ago707 ptsView detailsJoin discussion
Muse – Meta’s personal AI agent
ai.meta.com9 days ago658 ptsView detailsJoin discussion
AlphaGenome Atlas: a high-resolution map of human DNA
blog.google9 days ago593 ptsView detailsJoin discussion
I-have-ADHD: A skill to stop coding agents from burying the answer
github.com9 days ago538 ptsView detailsJoin discussion
Tao: Open math problems being non-renewably mined by AI
mathstodon.xyz9 days ago476 ptsView detailsJoin discussion
We Must Return to the Office to Use AI in Person
mcsweeneys.net9 days ago375 ptsView detailsJoin discussion
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
github.com9 days ago276 ptsView detailsJoin discussion
inceptionlabs.ai9 days ago245 ptsView detailsJoin discussion
Show HN: LLM Attention Visualization
ishamf.dev9 days ago176 ptsView detailsJoin discussion
Flights cancelled at UK airports due to ATC issue
bbc.com9 days ago126 ptsView detailsJoin discussion
27.5KB language-agnostic WebGPU syntax highlighter
gpu-lexer.vercel.app9 days ago119 ptsView detailsJoin discussion
Python sets and dictionaries can have quadratic-time performance
lemire.me9 days ago112 ptsView detailsJoin discussion
Multi-Agents LLM Financial Trading Framework
github.com10 days ago120 ptsView detailsJoin discussion
AI Responsibility – OpenAI and Anthropic
twitter.com9 days ago89 ptsView detailsJoin discussion
mhacevedo.com9 days ago68 ptsView detailsJoin discussion
DOJ Blocked ICE Agent Shooting Charge over Federal Prosecutor's Objections
propublica.org9 days ago55 ptsView detailsJoin discussion
grapheneos.social9 days ago52 ptsView detailsJoin discussion
How Climate Resilient Are the Largest Cities?
alphageo.ai9 days ago50 ptsView detailsJoin discussion
Meta Failed to Catch Hundreds of AI Child Abuse Ads
wired.com9 days ago44 ptsView detailsJoin discussion
Function Arguments Are Not Function Colors
jerf.org9 days ago44 ptsView detailsJoin discussion
"Please Remove All Mannered Prose" and Other LLM Incantations
matthewritch.com9 days ago37 ptsView detailsJoin discussion
AI Giants Work Hand-in-Hand with The Pentagon, Contracts Reveal
theintercept.com9 days ago34 ptsView detailsJoin discussion
Cognition (Devin) raises $2B at $48B valuation
cognition.com9 days ago23 ptsView detailsJoin discussion
After Work, We'll Have Each Other
asteriskmag.com9 days ago22 ptsView detailsJoin discussion
UAE-based Falcon AI NSFW classifier among top global open-source models (2025)
middleeastainews.com10 days ago23 ptsView detailsJoin discussion
Codex silently begs agents to make arbitrary web requests
spader.zone9 days ago19 ptsView detailsJoin discussion
Chinese AI Companies Conducting Distillation Campaigns Against U.S. AI Companies [pdf]
media.defense.gov9 days ago19 ptsView detailsJoin discussion
Has OpenAI model solved 80-year-old Navier-Stokes problem?
firstpost.com9 days ago20 ptsView detailsJoin discussion
10x More Efficient Pretraining
magic.dev9 days ago18 ptsView detailsJoin discussion
Why AI-generated slopaganda works
machinesociety.ai9 days ago18 ptsView detailsJoin discussion
Electrostatic Cathode Ray Tube Project 1 (2014)
labguysworld.com10 days ago19 ptsView detailsJoin discussion
Anthropic Researcher Quits over 'Out-of-Control' AI Fears
wsj.com9 days ago16 ptsView detailsJoin discussion
Show HN: VolAnti – Open-source acoustic detector for fibre-optic FPV drones
github.com9 days ago16 ptsView detailsJoin discussion
Mamdani: Officials 'lied' to New Yorkers about still-toxic after 9/11
independent.co.uk9 days ago15 ptsView detailsJoin discussion
New Yorkers were misled about air quality after 9/11
nytimes.com9 days ago14 ptsView detailsJoin discussion
dominis.blog9 days ago14 ptsView detailsJoin discussion
A Doctor Sued 700 Patients for Debts; 81 Were Arrested. Now He's a Senator
nytimes.com9 days ago13 ptsView detailsJoin discussion
Netanyahu Received an Explicit Warning Days Before Oct. 7
haaretz.com10 days ago14 ptsView detailsJoin discussion
Super Smash Brothers Melee has been 100% decompilated with the help of LLMs
github.com9 days ago12 ptsView detailsJoin discussion
Using radiative coatings to cool buildings on the cheap
economist.com9 days ago11 ptsView detailsJoin discussion
Show HN: Sparrow-2 – Noise cancellation isn't designed for conversational AI
sparrow2.tavuslabs.org9 days ago11 ptsView detailsJoin discussion
mastodon.social10 days ago11 ptsView detailsJoin discussion
Magic-mushroom compound blocks a side effect of chemotherapy in mice
nature.com10 days ago11 ptsView detailsJoin discussion
AI may have just solved a million-dollar math problem
scientificamerican.com9 days ago10 ptsView detailsJoin discussion
- Primary source
How GPT-5.6 Sol helps run quantum computing experiments
See how an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.
openai.com9 days agoView details
- Primary source
Explore how more capable, affordable AI can expand the work people and businesses can accomplish—and make growth more economical.
openai.com9 days agoView details
- Primary source
Introducing ChatGPT Images 2.5
ChatGPT Images 2.5 helps turn your ideas, sketches, and reference photos into more personalized, polished images that better reflect your ideas.
openai.com9 days agoView details
- Primary source
On the Navier–Stokes Millennium Prize Problem
We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.
openai.com9 days agoView details
- Primary source
Funding grants for new research into AI and teen development
Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.
openai.com10 days agoView details
- Primary source
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
huggingface.co9 days agoView details
- Primary source
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
AlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.
deepmind.google9 days agoView details
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
Today, Meta has introduced Muse, a personal AI agent that takes actions rather than just answering questions. Muse can send emails, book travel, negotiate bills, and pursue long term goals. It keeps working after you close the app and returns only when it needs approval. The bigger story for AI devs is architectural.…
marktechpost.com9 days agoView details
NVIDIA has announced CUDA Rust, a push to make Rust a first-class language for GPU kernels. Two open-source NVlabs projects cover the two CUDA programming models: cuda-oxide compiles SIMT kernels from Rust MIR through Pliron and LLVM to PTX, while cutile-rs JIT-compiles Tile kernels through CUDA Tile IR on stable Rust…
marktechpost.com9 days agoView details
DeepMind's AlphaGenome Atlas maps every single-letter change in the human genome with 1 impact score per variant. The post Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants appeared first on MarkTechPost.
marktechpost.com9 days agoView details
Drama swirls around OpenAI’s legendary mathematical milestone
OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired. In a blog post on Tuesday, OpenAI announced that it discovered a solution to the Navier-Stokes problem - which relates to the flow of liquid and gas - using an in…
theverge.com9 days agoView details
ChatGPT Sketch turns your bad drawings into detailed AI images
I used ChatGPT and its Sketch tool to make this AI-generated image of a cat. OpenAI announced ChatGPT Images 2.5 on Tuesday and is adding a new way to tell ChatGPT what you want it to make an image of: by drawing a doodle. With a new feature called Sketch, you can just draw something right inside ChatGPT and then tell…
theverge.com9 days agoView details
Muse, Meta’s New Personal AI Agent, Needs You to Trust It
Designed to compete with OpenClaw and Instinct, the company says Muse can do everything from sell your car to book you a plane ticket.
wired.com9 days agoView details
Meta bets on AI agent Muse to catch up in AI race
Meta is making another push to bring artificial intelligence to the masses with Muse, a personal assistant it says can put AI in the hands of virtually anyone. The product is the latest step in a multi-billion-dollar strategy overhaul designed to revitalize the company's ailing position in the AI race and help it catc…
theverge.com9 days agoView details
AI power users claim Anthropic duped them with subscriptions, and they’re taking it to court
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class actio…
theverge.com9 days agoView details
OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul
A landmark announcement by the frontier AI lab has been overshadowed by accusations of impropriety.
wired.com9 days agoView details
Google’s Atlas of the human genome could pave the way for new treatments
Google DeepMind has unveiled an AI tool that its scientists claim could help unravel the mysteries of the human genome and transform our understanding of biology, accelerating scientific research and ultimately paving the way for new treatments for diseases. The platform, called AlphaGenome Atlas, contains a "predicti…
theverge.com9 days agoView details
Adobe is trying to make its AI generators idiot-proof in Premiere
Suspenseful clock ticking… as you wait for AI to take your job. | Image: Adobe Adobe is overhauling how editors interact with AI in its Premiere professional video editing software. Its new Generative Media tool makes it easier to generate video, sound effects, music, and soundscapes without ever leaving the project t…
theverge.com9 days agoView details
CriticGen: Generation-Aware Evaluation as Actionable Feedback
arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framewo…
arxiv.org9 days agoView details
arXiv:2609.06815v1 Announce Type: new Abstract: An open, networked web will allow agents to run frozen models from multiple vendors, keep their history private, and teach each other which tool to call and when. Flat text (prompts, example pools) makes it difficult for the protocol to distinguish between noise statisti…
arxiv.org9 days agoView details
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems
arXiv:2609.05928v1 Announce Type: new Abstract: Large language models now compute correct tax liabilities on over 90% of well-formed cases in statutory benchmarks, which makes them candidates for the tax-advisory and compliance systems that consume such an answer directly. Real legal inputs, however, are frequently de…
arxiv.org9 days agoView details
Some Tokens Behave like Magnets: Revealing Linguistic Organization in the Layers of Language Models
arXiv:2609.05743v1 Announce Type: new Abstract: We identify a special group of token vectors inside large language models (LLMs), which we term magnetic vectors, that organize the surrounding tokens by either attracting or repelling them. Particularly, tokens pointing the same way as an attracting magnet are elongated…
arxiv.org9 days agoView details
AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable in…
arxiv.org9 days agoView details
Better Together: Complementary Query Rewriting Under a Strong RAG Baseline
arXiv:2609.05637v1 Announce Type: new Abstract: A popular way to improve Retrieval-Augmented Generation (RAG) is to rewrite the user's question into several variants and search with all of them. We test whether this actually helps once the underlying search is already strong. Under one fixed, competitive pipeline (BGE…
arxiv.org9 days agoView details
TamilEOT: A Dataset and Model for Semantic End-of-Turn Detection in Tamil Telephone Speech
arXiv:2609.05631v1 Announce Type: new Abstract: A voice agent has to decide, at every pause, whether the user has finished speaking. Without a model of the language that decision falls back to a fixed silence timeout: set it short and the agent interrupts, set it long and every turn pays the full wait. Open semantic e…
arxiv.org9 days agoView details
PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories
arXiv:2609.05488v1 Announce Type: new Abstract: Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic label. We introduce PGP-Clinical-TimeKAN, a trajectory-first framework for joint probabilistic forecasting of multivariate physiology. It combines missingness-aware t…
arxiv.org9 days agoView details
Deep belief networks are exact
arXiv:2609.05572v1 Announce Type: new Abstract: We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite parameters. This answers a question of Sutskever and Hinton. The proof upgrades their probability-sharing approximation to exac…
arxiv.org9 days agoView details
Discovering Translation-Worthy Languages with E-Values
arXiv:2609.06593v1 Announce Type: new Abstract: Choosing when to translate multilingual documents is a central routing problem in text classification: translation can improve predictions for some languages while degrading others or adding unnecessary computation. Uniform translation and heuristic language tiers do not…
arxiv.org9 days agoView details
Exposing Weaknesses in Emotion Recognition in Conversations
arXiv:2609.05806v1 Announce Type: new Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies.…
arxiv.org9 days agoView details
Generator-Independent Runtime Assurance under Partial Observation
arXiv:2609.06036v1 Announce Type: new Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtime verification gates. We ask when the closed-loop safety guarantee decouples from the generator. The prevailing per-cand…
arxiv.org9 days agoView details
You Are What You Read: Misalignment via In-Context Persona Induction
arXiv:2609.06851v1 Announce Type: new Abstract: Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in…
arxiv.org9 days agoView details
Compiling VGDL into Causal Models
arXiv:2609.05459v1 Announce Type: new Abstract: Reinforcement learning and large language models often struggle to accurately capture the causal mechanics of game environments. Standard reinforcement learning agents tend to rely on spurious correlations, while large language models are prone to hallucinating game rule…
arxiv.org9 days agoView details
A Rubric-Guided Large Language Model Solution for Opioid Use Disorder Computable Phenotyping
arXiv:2609.05682v1 Announce Type: new Abstract: Opioid use disorder (OUD) remains a public health crisis in the United States, yet it is difficult to identify from electronic health records (EHRs) because missing diagnosis codes and supporting evidence are buried in clinical narratives. Accurate OUD identification is…
arxiv.org9 days agoView details
Intra-Prompt Parallel Decoding for Common-Context Question Answering
arXiv:2609.05707v1 Announce Type: new Abstract: In common-context question answering (CCQA) tasks, multiple input questions share a common context to base their answers from. However, Large Language Models typically generate each answer using an independent prompt. While existing batching and caching techniques help i…
arxiv.org9 days agoView details
arXiv:2609.05757v1 Announce Type: new Abstract: Identifying the target of emotional words or phrases in crisis situations, especially health-related ones, is important for understanding public concerns across cultural and linguistic contexts. We propose CrisisKD, a five-stage teacher--student knowledge distillation fr…
arxiv.org9 days agoView details
Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance
arXiv:2609.05797v1 Announce Type: new Abstract: Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the diffe…
arxiv.org9 days agoView details
Dynamic Lagging for Simultaneous Translation
arXiv:2609.05799v1 Announce Type: new Abstract: In cascaded simultaneous speech translation, the machine translation (MT) system cannot control the read--write schedule of the upstream recognizer: it must decide, from a growing source prefix, how much target text to commit. We make a sentence-trained, decoder-only LLM…
arxiv.org9 days agoView details
Beyond Cross-Lingual Transfer: Benchmarking Propagation Boundaries in Multilingual LLM Unlearning
arXiv:2609.05976v1 Announce Type: new Abstract: Large Language Model (LLM) unlearning aims to suppress target knowledge while preserving general capabilities. In multilingual settings, unlearning must additionally propagate within its intended linguistic scope. However, existing evaluations mainly measure cross-lingua…
arxiv.org9 days agoView details
CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
arXiv:2609.05708v1 Announce Type: new Abstract: Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opi…
arxiv.org9 days agoView details
Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation
arXiv:2609.05993v1 Announce Type: new Abstract: Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasi…
arxiv.org9 days agoView details
arXiv:2609.06000v1 Announce Type: new Abstract: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet. The real-valued hidden states of a standard Transformer are retained, while only an auxiliary phase channel is quantized…
arxiv.org9 days agoView details
arXiv:2609.06011v1 Announce Type: new Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a p…
arxiv.org9 days agoView details
Generating Adversarial Texts for Machine Translation via GRPO
arXiv:2609.06048v1 Announce Type: new Abstract: As machine translation (MT) systems continue to improve, standard benchmarks become less informative for exposing remaining weaknesses. Traditional methods for creating challenging test sets rely on expensive manual creation or curation, while automated approaches strugg…
arxiv.org9 days agoView details
EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent
arXiv:2609.05576v1 Announce Type: new Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its…
arxiv.org9 days agoView details
arXiv:2609.05663v1 Announce Type: new Abstract: We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, F…
arxiv.org9 days agoView details
DPH Parser: A Bottom-Up Grammar-Driven Parser for Joint Constituency and Dependency Analysis
arXiv:2609.06070v1 Announce Type: new Abstract: This paper presents Dependency-Phrase Hierarchy Parser (DPH Parser), a grammar-driven bottom-up unsupervized parsing framework inspired by Generalized Phrase Structure Grammar (GPSG) and Head-driven Phrase Structure Grammar (HPSG). The parser incrementally constructs con…
arxiv.org9 days agoView details
From Two Passes to One: Compact and Efficient Target-Stance Extraction
arXiv:2609.06108v1 Announce Type: new Abstract: Target-Stance Extraction (TSE) is the task of predicting both the target (or topic) of an author's writing and the author's stance toward it. Existing approaches to TSE use a sequential pipeline of two separate neural models: one to identify the target and another to det…
arxiv.org9 days agoView details
STQA: A Benchmark for Stock-Focused Tabular Question Answering over Historical and Forecasted Data
arXiv:2609.06117v1 Announce Type: new Abstract: Stock market analysis inherently requires composite reasoning over historical records and future projections, yet existing benchmarks remain fragmented across isolated tasks. We introduce STQA (Stock-focused Tabular Question Answering), an end-to-end benchmark designed t…
arxiv.org9 days agoView details
Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement
arXiv:2609.06189v1 Announce Type: new Abstract: This study examines how Customer Relationship Management (CRM), Master Data Management (MDM), and Customer Knowledge Management (CKM) jointly constitute a Customer Relationship Intelligence (CRI) framework for enhanced Customer Engagement (CE). A cross-sectional survey o…
arxiv.org9 days agoView details
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
arXiv:2609.06263v1 Announce Type: new Abstract: Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emergi…
arxiv.org9 days agoView details
Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin
arXiv:2609.06266v1 Announce Type: new Abstract: Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The…
arxiv.org9 days agoView details
AutoKD: Autonomous Knowledge Discovery
arXiv:2609.06366v1 Announce Type: new Abstract: Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun…
arxiv.org9 days agoView details
A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation
arXiv:2609.06690v1 Announce Type: new Abstract: Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, compu…
arxiv.org9 days agoView details
SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use
arXiv:2609.06124v1 Announce Type: new Abstract: High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are critical to long-horizon tool use. As a result, even when a model selects the correct tool, t…
arxiv.org9 days agoView details
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
arXiv:2609.05837v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limite…
arxiv.org9 days agoView details
What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction
arXiv:2609.05882v1 Announce Type: new Abstract: Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how t…
arxiv.org9 days agoView details
arXiv:2609.05843v1 Announce Type: new Abstract: Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs…
arxiv.org9 days agoView details
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
arXiv:2609.06289v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically valid…
arxiv.org9 days agoView details
LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
arXiv:2609.06079v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic representations that evolve across layers, from local visual geometry to abstract, language-ali…
arxiv.org9 days agoView details
The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble
arXiv:2609.05894v1 Announce Type: new Abstract: The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era that has been driven by the growth in large language models (LLMs), large-scale compute infrastructures, and autonomous reasoning systems. However, the rapid accelerati…
arxiv.org9 days agoView details
Decomposing LLM-Judge Uncertainty to Target Expert Labels
arXiv:2609.06444v1 Announce Type: new Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which…
arxiv.org9 days agoView details
DFlow: Enabling Verifier Information Flow in Block Diffusion Speculative Decoding
arXiv:2609.06498v1 Announce Type: new Abstract: Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the reje…
arxiv.org9 days agoView details
Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring
arXiv:2609.06315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value…
arxiv.org9 days agoView details
arXiv:2609.06527v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL…
arxiv.org9 days agoView details
UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms
arXiv:2609.05910v1 Announce Type: new Abstract: Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack…
arxiv.org9 days agoView details
arXiv:2609.05553v1 Announce Type: new Abstract: Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before…
arxiv.org9 days agoView details
SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives
arXiv:2609.06611v1 Announce Type: new Abstract: Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates…
arxiv.org9 days agoView details
arXiv:2609.06324v1 Announce Type: new Abstract: Autoregressive language models commit one token per forward pass; diffusion language models commit a block of tokens over several steps. We ask whether a block can be committed in a single forward pass. We study this with a noise-conditioned masked denoiser: a data-indep…
arxiv.org9 days agoView details
Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space
arXiv:2609.06381v1 Announce Type: new Abstract: Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languages, enabling cross-lingual transfer to improve multilingual capabilities. Previous CLA methods often ignore language-specific information encoded in representati…
arxiv.org9 days agoView details
Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus
arXiv:2609.06634v1 Announce Type: new Abstract: LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they are trained upon. To examine if Multiword Expressions (MWEs) still set a bottleneck for LLMs regarding language understand…
arxiv.org9 days agoView details
arXiv:2609.06406v1 Announce Type: new Abstract: Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate o…
arxiv.org9 days agoView details
InsightChain: Optimized Chain-of-Insight Analytics for LLM-driven Data Visualization
arXiv:2609.06438v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches often frame visualization generation as a single-step mapping from user query to figure or code, overlooking the iterative analytical reasoning process of expert…
arxiv.org9 days agoView details
arXiv:2609.06545v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associ…
arxiv.org9 days agoView details
ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics
arXiv:2609.06663v1 Announce Type: new Abstract: Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead of KV cache storage. Recent KV cache eviction approaches incorporate a cosine similarity-based diversity metric with impor…
arxiv.org9 days agoView details
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
arXiv:2609.06702v1 Announce Type: new Abstract: Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to…
arxiv.org9 days agoView details
arXiv:2609.06703v1 Announce Type: new Abstract: High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-graine…
arxiv.org9 days agoView details
arXiv:2609.06771v1 Announce Type: new Abstract: Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering…
arxiv.org9 days agoView details
arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that bench…
arxiv.org9 days agoView details
The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
arXiv:2609.05514v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-gro…
arxiv.org9 days agoView details
Factors Influencing the Emergence of Dependency Length Minimization in Neural Agent Simulations
arXiv:2609.06025v1 Announce Type: new Abstract: Given various grammatical options, language users prefer the word order choice that reduces the overall length of syntactic dependencies, a principle known as dependency length minimization (DLM). The origins of this preference remain an open question, particularly wheth…
arxiv.org9 days agoView details
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will en…
arxiv.org9 days agoView details
arXiv:2609.05531v1 Announce Type: new Abstract: No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO…
arxiv.org9 days agoView details
Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools return reliable information. However, tool returns in real-world systems can be plausible yet in…
arxiv.org9 days agoView details
From Narrative to Auditable Forecasts: A Structured Scaffold for Agentic Forecasting
arXiv:2609.05905v1 Announce Type: new Abstract: LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events. However, current agentic forecasting often relies on implicit narrative aggregation: agents collect evidence, discuss it…
arxiv.org9 days agoView details
ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models
arXiv:2609.05461v1 Announce Type: new Abstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space: an action is preferred if its predicted future embedding lands closer to the goal embedding. This silently assumes that latent closeness is action-rankable, i.e., t…
arxiv.org9 days agoView details
Distilling Vision-Language Models for On-Device Fire Understanding
arXiv:2609.05782v1 Announce Type: new Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this…
arxiv.org9 days agoView details
Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
arXiv:2609.05736v1 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool…
arxiv.org9 days agoView details
Inference-Time Graph Engineering for Multi-Agent LLM Workflows
arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temp…
arxiv.org9 days agoView details
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
arXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline…
arxiv.org9 days agoView details
RAPID: Reliability-Aware Pair Importance Distillation
arXiv:2609.05481v1 Announce Type: new Abstract: Inter example relational distillation transfers a teacher's representation geometry by matching relations among examples within a mini batch. Computing all pairs has quadratic complexity in the batch size, whereas uniform subsampling may use a limited relation budget ine…
arxiv.org9 days agoView details
Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
arXiv:2609.05800v1 Announce Type: new Abstract: Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dim…
arxiv.org9 days agoView details
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening…
arxiv.org9 days agoView details
SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
arXiv:2609.05511v1 Announce Type: new Abstract: Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step,…
arxiv.org9 days agoView details
arXiv:2609.05527v1 Announce Type: new Abstract: Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it…
arxiv.org9 days agoView details
arXiv:2609.05643v1 Announce Type: new Abstract: This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia, national laboratories, and industry who are reshaping materials science discovery. The meeting explored how AI, autonomous agents, self-driving labs, higher performance and…
arxiv.org9 days agoView details
arXiv:2609.05758v1 Announce Type: new Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model path, or separate those roles. We report a production migration of a customer-support assistant at a large accommodation marketplace (millions of conversations per…
arxiv.org9 days agoView details
More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
arXiv:2609.05788v1 Announce Type: new Abstract: Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We eval…
arxiv.org9 days agoView details
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
arXiv:2609.05824v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers ty…
arxiv.org9 days agoView details
Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
arXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain conf…
arxiv.org9 days agoView details
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review
arXiv:2609.05947v1 Announce Type: new Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisio…
arxiv.org9 days agoView details
MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search
arXiv:2609.05992v1 Announce Type: new Abstract: As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable agent must not only achieve high task completion accuracy but also perform well in interaction quality, safety, and efficiency, raising a central question: can these…
arxiv.org9 days agoView details
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
arXiv:2609.06059v1 Announce Type: new Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are oft…
arxiv.org9 days agoView details
Recovering Temporal and Geographic Signals from Language Model Embeddings
arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a simple projection-based method that operat…
arxiv.org9 days agoView details
Explaining AI Agents Through Execution Traces
arXiv:2609.06063v1 Announce Type: new Abstract: AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, tra…
arxiv.org9 days agoView details
Generating Instance Generators in PDDL Planning
arXiv:2609.06071v1 Announce Type: new Abstract: PDDL, the de-facto standard language in the AI Planning community, is designed to specify planning domains: sets of instances that share the same predicates and action schemas. Yet it does not provide any means to specify the actual instance set, i.e., legality constrain…
arxiv.org9 days agoView details
CWF: A Collaborative Writing Framework for Personalized and Reliable Popular Science Writing
arXiv:2609.06126v1 Announce Type: new Abstract: We introduce Personalized and Reliable Popular Science Writing, a novel task that requires adapting scientific explanations to audiences with different cognitive levels while preserving factual accuracy. However, improving personalization often introduces simplifications…
arxiv.org9 days agoView details
Substrate-Portable Execution for Production LLM Workflows
arXiv:2609.06128v1 Announce Type: new Abstract: Production LLM agents execute tool-calling loops, retrieval chains, and compositional workflows in multiple modes, yet execution semantics are often coupled to one runtime. We encountered this portability problem in Rufus, a conversational AI assistant with a large tool…
arxiv.org9 days agoView details
IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing
arXiv:2609.06131v1 Announce Type: new Abstract: Environmental identification in wireless sensing is essential for 6G integrated sensing and communication (ISAC) systems to achieve reliable situational awareness. However, deep learning (DL) models for this task often fail to generalize under domain shift across diverse…
arxiv.org9 days agoView details
SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores
arXiv:2609.06192v1 Announce Type: new Abstract: Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agen…
arxiv.org9 days agoView details
Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts
arXiv:2609.06194v1 Announce Type: new Abstract: Offshore wind turbines are widely used to generate renewable energy, but their maintenance can result in decreased efficiency due to forced shutdowns. Accurate wind turbine power predictions can identify periods of low power that would be ideal for scheduling maintenance…
arxiv.org9 days agoView details
arXiv:2609.06391v1 Announce Type: new Abstract: Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can plan retrieval, traverse relations, verify intermediate claims, delegate subtasks, and use tools. This combination is useful when answers depend on relations acro…
arxiv.org9 days agoView details
From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
arXiv:2609.06403v1 Announce Type: new Abstract: Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout gra…
arxiv.org9 days agoView details
Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification
arXiv:2609.06445v1 Announce Type: new Abstract: A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditio…
arxiv.org9 days agoView details
When and What to Teach: Budget-Aware Online Adaptation for Web Agents
arXiv:2609.05513v1 Announce Type: new Abstract: Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners mu…
arxiv.org9 days agoView details
Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration
arXiv:2609.05801v1 Announce Type: new Abstract: A document modeled as a discrete sequence of tokens can be thought of as being generated from a composition of texts from different domains; a README file, for example, moves between prose, code, and configuration. When such a document is corrupted and only frozen domain…
arxiv.org9 days agoView details
arXiv:2609.05821v1 Announce Type: new Abstract: Vision-language models (VLMs) often answer new questions about recurring visual content, where reusing the key-value (KV) cache can avoid re-encoding expensive visual prefixes. Exact-prefix reuse, however, fails when the same visual content appears under a changed prefix…
arxiv.org9 days agoView details
Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which rais…
arxiv.org9 days agoView details
AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents
arXiv:2609.05802v1 Announce Type: new Abstract: Large language models answering questions over multi-page documents are expected to cite the supporting pages, yet supplied citations are sometimes inaccurate, and current evaluations score citations at generation time or against text passages: no existing benchmark eval…
arxiv.org9 days agoView details
XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
arXiv:2609.06842v1 Announce Type: new Abstract: When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., "How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception ("regex are fragile") and meaningfully direct the…
arxiv.org9 days agoView details
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
arXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence…
arxiv.org9 days agoView details
AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection
arXiv:2609.05899v1 Announce Type: new Abstract: Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit mod…
arxiv.org9 days agoView details
Neuron-Guided Fine-Tuning: Unlocking Efficient Alignment Mechanisms for Large Language Models
arXiv:2609.05913v1 Announce Type: new Abstract: Existing Supervised Fine-Tuning paradigms, particularly Full Parameter Fine-Tuning are often plagued by parameter redundancy, inconsistent data quality, and catastrophic forgetting, which current methods typically address in isolation and lack a unified optimization sign…
arxiv.org9 days agoView details
arXiv:2609.05938v1 Announce Type: new Abstract: Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to…
arxiv.org9 days agoView details
The Blindness of Document-Level Translation Evaluation
arXiv:2609.05949v1 Announce Type: new Abstract: Document-level machine translation (MT) evaluation extends segment-level protocols by presenting full documents to annotators, on the assumption that such presentation elicits document-level judgments. We test this assumption with a counterfactual condition (MIX) in whic…
arxiv.org9 days agoView details
Don't Lose Entities from Retrieval to Generation: Dual Entity Recovery RAG for multi-hop QA
arXiv:2609.06065v1 Announce Type: new Abstract: Retrieval-augmented multi-hop question answering (QA) decomposes a query into sub-questions and decomposes the corpus into smaller retrieval units such as sentences. Both forms of decomposition improve the pipeline, but we show that both share the same vulnerability, the…
arxiv.org9 days agoView details
arXiv:2609.06129v1 Announce Type: new Abstract: Agents that talk across organizations exchange long messages billed by the token. A shorter notation therefore looks like a saving that costs nothing but an agreement to use it. Recent work reports the saving is conditional. Compressed notation can instead raise total to…
arxiv.org9 days agoView details
arXiv:2609.06188v1 Announce Type: new Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic, and visual modalities. Although large language models (LLMs) offer powerful language understanding, adapting them to m…
arxiv.org9 days agoView details
What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark
arXiv:2609.06147v1 Announce Type: new Abstract: Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited ou…
arxiv.org9 days agoView details
arXiv:2609.06212v1 Announce Type: new Abstract: LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Ef…
arxiv.org9 days agoView details
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for…
arxiv.org9 days agoView details
Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
arXiv:2609.05677v1 Announce Type: new Abstract: Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded…
arxiv.org9 days agoView details
Agentic Pressure: The Endogenous Entropy of Reliable Autonomy
arXiv:2609.05995v1 Announce Type: new Abstract: Achieving reliable autonomy in the wild requires agents to sustain continuous operations across long-horizon trajectories. However, as agents navigate these unconstrained settings, they encounter cumulative friction that inherently destabilizes their alignment. In this p…
arxiv.org9 days agoView details
MedWER: A Reproducible, Model-Free Evaluation Protocol for Medical Speech Recognition
arXiv:2609.05728v1 Announce Type: new Abstract: Overall word error rate hides clinically critical errors: a transcript can be 95% correct and still swap one drug for another. The usual fix weights errors on medical entities, and almost always depends on an evaluation-time named-entity recognition (NER) model or cloud…
arxiv.org9 days agoView details
arXiv:2609.06410v1 Announce Type: new Abstract: Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failing to capture temporal cues, multi-angle views and fine-grained visual details. Directly applying video vision-language models (VLMs) to product AVE results in li…
arxiv.org9 days agoView details
The Normalization of Deviance in AI Development
arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the inst…
arxiv.org9 days agoView details
Planning and Scheduling Business Processes under Control-Flow Uncertainty
arXiv:2609.05578v1 Announce Type: new Abstract: Scheduling activities in business processes can improve efficiency (e.g., reduce makespan), but is challenging because the exact sequence of activities required to complete a case is often uncertain due to decisions based on data that emerges during execution. Neverthele…
arxiv.org9 days agoView details
Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction
arXiv:2609.06731v1 Announce Type: new Abstract: Temporal relation extraction determines whether an event occurs before, after, or simultaneously with another event, and therefore relies on accurately modeling how the two events interact. Mainstream systems achieve this by concatenating event spans or using shallow fus…
arxiv.org9 days agoView details
Damage-Aware Bandit Pruning for Vision and Language Transformers
arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fi…
arxiv.org9 days agoView details
Models
View all models →text-generation · transformers · safetensors · qwen3_5_moe
huggingface.co9 days ago824 ptsView details
text-generation · transformers · safetensors · qwen3_5_moe
huggingface.co9 days ago660 ptsView details
gguf · arxiv:2605.22064 · base_model:tencent/Hy-MT2-1.8B
huggingface.co10 days ago169 ptsView details
sentence-similarity · sentence-transformers · safetensors · qwen3_vl
huggingface.co10 days ago8 ptsView details
text-generation · safetensors · qwen3 · router
huggingface.co10 days ago1 ptsView details
batuhan-elibuyuk/logicer-ggufs-public
safetensors · gguf · endpoints_compatible
huggingface.co10 days agoView details
Open source
View all open source →## [3.10.0](https://github.com/openai/openai-python/compare/v3.9.0...v3.10.0) (2026-09-08) ### Features * **api:** add GPT Image 2.5 models and image options ([#3824](https://github.com/openai/openai-python/issues/3824)) ([5b39c45](https://github.com/openai/openai-python/commit/…
github.com9 days agoView details
## [3.9.0](https://github.com/openai/openai-python/compare/v3.8.0...v3.9.0) (2026-09-05) ### Features * **api:** Add prompt cache diagnostics ([#3800](https://github.com/openai/openai-python/issues/3800)) ([8326784](https://github.com/openai/openai-python/commit/83267847a0219ea8…
github.com9 days agoView details
langchain-ai/langchain langchain-openai==1.6.1
Changes since langchain-openai==1.6.0 fix(openai): bump `max_completion_tokens` in cache breakpoint integration test (#40284) release(openai): 1.6.1 (#40268) chore(model-profiles): refresh model profile data (#40217) fix(openai): support Azure AD auth with OpenAI 3.8 (#40190) fe…
github.com9 days agoView details