Skip to content

Archive / 2026-09-14

September 14, 2026

  1. The AI job market in 2026

    ilinmaks.com3 days ago65 ptsView detailsJoin discussion

  2. Graphic Rants: Nanite Tessellation

    graphicrants.blogspot.com3 days ago26 ptsView detailsJoin discussion

  3. AI Is in Dangerous Hands

    wheresyoured.at3 days ago17 ptsView detailsJoin discussion

  4. The Register of British Slave-Traders

    britishslavetraders.org3 days ago10 ptsView detailsJoin discussion

  5. Primary source

    How Fyxer built an AI executive assistant people trust

    Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.

    openai.com3 days agoView details

  6. Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second

    Meta engineering team introduced ZGateway, a proxy tier that now sits between client applications and ZippyDB, the Meta’s most widely used key value store. ZippyDB backs product metadata, counters, and configuration at billions of operations per second. ZGateway started as a fix for connection sprawl across more than…

    marktechpost.com3 days agoView details

  7. Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

    Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestr…

    marktechpost.com3 days agoView details

  8. Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

    A practitioner's map of the 3 layers in a modern agent stack, with verified sources and an overlap analysis. The post Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery appeared first on MarkTechPost.

    marktechpost.com3 days agoView details

  9. Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data

    Reward AI has released OM-1 (Omnibody Model 1), a general-purpose manipulation policy trained entirely on human demonstrations captured with a 7-DoF wearable glove, with no teleoperation or on-robot data. The policy runs on industrial arms and humanoids at human speed, learns a new task from under 30 minutes of data,…

    marktechpost.com3 days agoView details

  10. Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks

    Sakana AI researchers Jeffrey Seely and Julian Gould introduce Augmented Lagrangian Predictive Coding (PC-ALM), a local-learning alternative to backpropagation. By attaching a Lagrange multiplier to each layer constraint, PC-ALM keeps predictive coding's layer-local updates while recovering exact backprop gradients in…

    marktechpost.com3 days agoView details

  11. NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

    NVIDIA has open-sourced OSMO, the Kubernetes-native workflow orchestrator it uses internally for Project GR00T, Isaac Lab, and Isaac Sim. OSMO lets robotics teams define training, simulation, and hardware-in-the-loop tasks in a single YAML file and routes each one to the right compute tier, from GB200 clusters to Jets…

    marktechpost.com4 days agoView details

  12. Is Big Tech’s AI slowdown a safety pact or a cartel?

    When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to "pace the frontier," signing on at l…

    theverge.com3 days agoView details

  13. What execs and politicians are saying about slowing down AI development

    Dario Amodei kicked off a flood of statements over the past few days about AI safety by publishing a long essay titled "We Must Pace the Frontier" detailing why AI development should be slowed down. Other AI leaders and politicians are speaking out in favor of or opposing his points, and we've compiled some of them he…

    theverge.com3 days agoView details

  14. Jensen Huang puts Trump on speakerphone onstage to announce robots won’t take over the world

    NVIDIA CEO Jensen Huang speaks during the G20 Innovation Ministerial in Chapel Hill, North Carolina, on September 2, 2026. (Photo by Matt RAMEY / AFP via Getty Images) | AFP via Getty Images Nvidia CEO Jensen Huang took a call from President Trump on Monday while onstage at the All-In Podcast's All-In Summit. It's not…

    theverge.com3 days agoView details

  15. New York Seizes a Dozen Celebrity Deepfake Websites

    In the biggest-ever legal action against harmful deepfake websites, the Manhattan District Attorney’s Office has seized 12 sites that collectively targeted around 1,200 victims.

    wired.com3 days agoView details

  16. Microsoft says ‘people matter more than AI’ following safety concerns

    Microsoft is publishing a 37-page "humanist AI code of conduct" today, amid growing safety concerns over AI model progress. Anthropic CEO Dario Amodei called for a coordinated slow down of AI development over the weekend, after researchers warned recently that AI model progress could outpace our ability to safely depl…

    theverge.com3 days agoView details

  17. AI Leaders Are Calling for a Slowdown. Trump’s Team Says It’s on Them

    Sam Altman and Elon Musk backed Anthropic CEO Dario Amodei’s weekend plea for regulation. The White House seems unlikely to oblige.

    wired.com3 days agoView details

  18. Sexually Explicit Deepfake Sites Target 100-Plus Politicians in Europe

    An analysis of 160 deepfake websites reveals politicians in 22 countries appear on them. Nearly all of them are women.

    wired.com3 days agoView details

  19. ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction

    While the publishing industry frets over how authors are using AI, many readers are taking things into their own hands.

    wired.com4 days agoView details

  20. From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

    arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog…

    arxiv.org3 days agoView details

  21. Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

    arXiv:2609.13556v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within thes…

    arxiv.org3 days agoView details

  22. Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

    arXiv:2609.13582v1 Announce Type: new Abstract: A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whe…

    arxiv.org3 days agoView details

  23. $\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents

    arXiv:2609.13602v1 Announce Type: new Abstract: Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. We introduce $\tau$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms…

    arxiv.org3 days agoView details

  24. Enhancing Event Candidate Acquisition for Event Linking

    arXiv:2609.13670v1 Announce Type: new Abstract: Event linking associates event mentions in text with entries in a knowledge base (KB), or identifies them as out-of-KB events. Although existing methods use different architectures, candidate event acquisition can still be weakened by short ambiguous mentions, noisy argu…

    arxiv.org3 days agoView details

  25. Recoverability as a System Primitive for Long-Horizon AI Agents

    arXiv:2609.13672v1 Announce Type: new Abstract: AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward. A saved state is not necessarily a suitable place t…

    arxiv.org3 days agoView details

  26. LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference

    arXiv:2609.13682v1 Announce Type: new Abstract: We introduce LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing (trained via a straight-through estimator) with joint LoRA fine-tuning. LayerRoute augments each of the 24 transformer blocks in Qwen2…

    arxiv.org3 days agoView details

  27. Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

    arXiv:2609.13685v1 Announce Type: new Abstract: Negation remains a longstanding challenge for both language models (LMs) and large language models (LLMs). Prior work mainly focuses on a small set of high-frequency single-word negation cues, such as not and never, with limited exploration of broader negation types and…

    arxiv.org3 days agoView details

  28. Positioning manuscripts in the scientific landscape with agentic AI

    arXiv:2609.13760v1 Announce Type: new Abstract: Publishing a research manuscript is a routine yet demanding part of scientific life: time-consuming, stressful, and often uncertain in outcome. Recent advances in large language model (LLM)-based agentic AI have shown promise across a range of scientific tasks, and here…

    arxiv.org3 days agoView details

  29. Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs

    arXiv:2609.13776v1 Announce Type: new Abstract: Integrating heterogeneous relational databases into a centralized ontology remains a persistent challenge in enterprise knowledge representation, primarily due to semantic heterogeneity, cryptic schema naming, missing metadata, and the abstraction gap between relational…

    arxiv.org3 days agoView details

  30. Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging

    arXiv:2609.13721v1 Announce Type: new Abstract: Quantum Natural Language Processing (QNLP) uses pregroup grammars to translate grammatical structure into diagrammatic representations and quantum circuits. Recent Hindi QNLP work has shown that Hindi-specific pregroup grammars can support grammar-sensitive compositional…

    arxiv.org3 days agoView details

  31. Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

    arXiv:2609.13745v1 Announce Type: new Abstract: Vision--language models (VLMs) can answer chart questions accurately, but output accuracy does not show how they combine the evidence needed to recover an exact value. We study vertical-bar value reading with controlled counterfactual activation patching in Qwen2.5VL-7B-…

    arxiv.org3 days agoView details

  32. HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering

    arXiv:2609.13768v1 Announce Type: new Abstract: Multi-hop question answering often fails when retrieval treats evidence as isolated matches to the original question, since the facts needed to answer a complex question are usually connected through intermediate entities, relations, and constraints. We propose HyperProv…

    arxiv.org3 days agoView details

  33. SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection

    arXiv:2609.13794v1 Announce Type: new Abstract: Detecting harmful memes is critical for maintaining safe online communities. However, harmful intent is often implicit, arising from visual-textual incongruity and cultural stereotypes, which challenges existing multimodal detectors. We propose SyRHM, a framework that de…

    arxiv.org3 days agoView details

  34. Understanding the Limits of Agentic ICD Coding

    arXiv:2609.13806v1 Announce Type: new Abstract: ICD-10-CM codes are alphanumeric codes used in the US to classify diagnoses and injuries for medical billing and epidemiological reporting. Standard ICD-10-CM benchmarks report aggregate metrics that obscure performance on complex coding scenarios. We evaluate neural, wo…

    arxiv.org3 days agoView details

  35. Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

    arXiv:2609.13841v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically examine this risk, we developed the Romanti…

    arxiv.org3 days agoView details

  36. Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+

    arXiv:2609.13847v1 Announce Type: new Abstract: In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja with Chichewa, and evaluate NLLB-200, Goo…

    arxiv.org3 days agoView details

  37. ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation

    arXiv:2609.13856v1 Announce Type: new Abstract: Enterprise customer support systems must answer customer questions correctly, retrieve the right policy information, use customer context, and pass difficult cases to human agents when needed. This paper presents ShopEase, a Generative AI-based multi-agent framework for…

    arxiv.org3 days agoView details

  38. SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity

    arXiv:2609.13874v1 Announce Type: new Abstract: Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible but cross-patient components, creating a s…

    arxiv.org3 days agoView details

  39. Phorecaster365: A Human-Supervised Reference Architecture for Hybrid Pharmaceutical Sales Forecasting and Planning Decision Support

    arXiv:2609.13907v1 Announce Type: new Abstract: Pharmaceutical sales forecasts inform planning across products, regions, and distribution channels, yet their interpretation depends on inventory availability, transaction semantics, product lifecycle, and the information available when each forecast is issued. A model p…

    arxiv.org3 days agoView details

  40. In the Blind: Building Pseudo-References for MT Evaluation

    arXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in whic…

    arxiv.org3 days agoView details

  41. North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)

    arXiv:2609.13916v1 Announce Type: new Abstract: We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same foundation as Cohere's Command A Plus, a mixture-of-experts architecture with 25 billion active parameters out of 218 bill…

    arxiv.org3 days agoView details

  42. Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models

    arXiv:2609.13534v1 Announce Type: new Abstract: We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises monotonically with transformer depth, whereas benign prompts remain flat or oscillatory. This cross-layer s…

    arxiv.org3 days agoView details

  43. Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus

    arXiv:2609.13936v1 Announce Type: new Abstract: Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_…

    arxiv.org3 days agoView details

  44. CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education

    arXiv:2609.13942v1 Announce Type: new Abstract: The CRITICS project addresses science accessibility and literacy by converging advanced Machine Translation (MT) based on Large Language Models (LLMs) with educational technology. By leveraging MT systems specifically optimized for scientific content, educational institu…

    arxiv.org3 days agoView details

  45. Thought without systematicity? Evaluating reasoning models on rule induction tasks

    arXiv:2609.13948v1 Announce Type: new Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance…

    arxiv.org3 days agoView details

  46. Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

    arXiv:2609.13997v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Yet problems beyond the model's current capability, where rollouts uniformly fail and no learning signal is produced, are…

    arxiv.org3 days agoView details

  47. Measuring the Creativity of Frontier LLMs in Automated Research

    arXiv:2609.14057v1 Announce Type: new Abstract: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Value…

    arxiv.org3 days agoView details

  48. GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning

    arXiv:2609.14066v1 Announce Type: new Abstract: Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness…

    arxiv.org3 days agoView details

  49. One Size Does Not Fit All: Setting Inference Depth from the Questions a Deployment Actually Asks

    arXiv:2609.14144v1 Announce Type: new Abstract: A transformer language model is trained to respond to any prompt, but each deployment asks only a narrow range of questions: a support assistant sees delivery complaints, a coding tool sees Python. Every deployment nonetheless pays the same computation per token. This pa…

    arxiv.org3 days agoView details

  50. Map Users and Mapmakers: The Scope of Cognitive Attribution from Acquired Representations

    arXiv:2609.13879v1 Announce Type: new Abstract: An acquired representation can enlarge a system's cognitive repertoire without transferring the capacities exercised in producing that representation. This paper develops a framework for specifying that enlargement and its limits. Its central contribution is a five-part…

    arxiv.org3 days agoView details

  51. When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

    arXiv:2609.14157v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a model handles questions that do not need t…

    arxiv.org3 days agoView details

  52. Towards Evolving Context Parameterization for Large Language Models

    arXiv:2609.14168v1 Announce Type: new Abstract: Context parameterization enables large language models (LLMs) to internalize contexts into reusable model parameters, avoiding repeated processing across subsequent queries. However, existing methods typically assume static contexts and lack explicit mechanisms for disti…

    arxiv.org3 days agoView details

  53. A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement

    arXiv:2609.14178v1 Announce Type: new Abstract: The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a pro…

    arxiv.org3 days agoView details

  54. Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had

    arXiv:2609.14121v1 Announce Type: new Abstract: The Semantic Web set out to give information a machine-interpretable form so that software could integrate and reason over it. Its standards became scientific knowledge infrastructure, but the machine competence it promised did not follow, and the systems now answering q…

    arxiv.org3 days agoView details

  55. LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects

    arXiv:2609.13918v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in specialized university courses, but control-systems questions require coordinated terminology, notation, derivations, and stepwise explanations. Direct general-purpose responses may be inconsistently structured and ha…

    arxiv.org3 days agoView details

  56. ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading

    arXiv:2609.13825v1 Announce Type: new Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) or limit-order-book depth features, leaving microstructure pattern theories from the practitioner literature, namely Auct…

    arxiv.org3 days agoView details

  57. The Attribution-Compression Frontier in Retrieval-Augmented Generation

    arXiv:2609.14245v1 Announce Type: new Abstract: Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does not characterize citation attribution. We measure citation attribution across compression methods and budgets, comparing reranking, extractive selection, abstract…

    arxiv.org3 days agoView details

  58. Document Topic Alignment Metrics for Evaluating Topic Models of Short-Text Public Health Communications on Social Media

    arXiv:2609.14256v1 Announce Type: new Abstract: Topic models are widely used to analyze public health-related social media short texts, yet their evaluation remains dominated by metrics that focus entirely on generated topics alone. There is a lack of metrics that quantitatively assess whether assigned topics meaningf…

    arxiv.org3 days agoView details

  59. DenMark: Robust Semantic Watermarking for Diffusion Language Models

    arXiv:2609.14257v1 Announce Type: new Abstract: Semantic text watermarks encode signals in meaning rather than surface token choices, offering robustness to paraphrasing and other semantic-preserving edits. Existing semantic watermarking methods are primarily designed for autoregressive language models (ARLMs), where…

    arxiv.org3 days agoView details

  60. Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

    arXiv:2609.13807v1 Announce Type: new Abstract: Large language models reason in high-dimensional hidden-state spaces, while users observe only final outputs. We introduce Bypass Observation, a non-intrusive layer-wise readout architecture that attaches read-only observation heads to selected Transformer layers without…

    arxiv.org3 days agoView details

  61. Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error

    arXiv:2609.14400v1 Announce Type: new Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory…

    arxiv.org3 days agoView details

  62. NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

    arXiv:2609.14448v1 Announce Type: new Abstract: Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated wit…

    arxiv.org3 days agoView details

  63. Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation

    arXiv:2609.14528v1 Announce Type: new Abstract: Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along paths in a KG to answer natural-language questions. However, existing KGQA systems typically focus on predicting the final answer without explicitly modeling or…

    arxiv.org3 days agoView details

  64. Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition

    arXiv:2609.14542v1 Announce Type: new Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541…

    arxiv.org3 days agoView details

  65. Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

    arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as…

    arxiv.org3 days agoView details

  66. Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

    arXiv:2609.13475v1 Announce Type: new Abstract: Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the…

    arxiv.org3 days agoView details

  67. Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management

    arXiv:2609.13552v1 Announce Type: new Abstract: Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation, and constraint checking. Although these tools can reduce workload and accelerate planning, their non-deterministic out…

    arxiv.org3 days agoView details

  68. Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

    arXiv:2609.13559v1 Announce Type: new Abstract: Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We add…

    arxiv.org3 days agoView details

  69. Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

    arXiv:2609.13637v1 Announce Type: new Abstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, be…

    arxiv.org3 days agoView details

  70. Solar Intelligence

    arXiv:2609.13648v1 Announce Type: new Abstract: Solar energy decision support is fragmented across dashboards that provide data without explanation, research papers are slow to parse, and general-purpose language models are not solar domain specific and answer without evidence. This paper introduces Solar Intelligence…

    arxiv.org3 days agoView details

  71. Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

    arXiv:2609.13680v1 Announce Type: new Abstract: Fine-tuning instruct models often improves target performance while inducing behavioral drift from the reference model, which can degrade existing capabilities. Rather than treating this drift as an uncontrolled consequence of optimization, we specify a behavioral drift…

    arxiv.org3 days agoView details

  72. JaxAHT: A JAX-Based Library for Ad Hoc Teamwork

    arXiv:2609.13716v1 Announce Type: new Abstract: Ad Hoc Teamwork (AHT) addresses the challenge of designing agents capable of coordinating with novel partners without prior coordination. However, progress in the field is hindered by the prohibitive computational cost of the AHT research lifecycle, the lack of standardi…

    arxiv.org3 days agoView details

  73. IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

    arXiv:2609.13725v1 Announce Type: new Abstract: An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived…

    arxiv.org3 days agoView details

  74. Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

    arXiv:2609.13731v1 Announce Type: new Abstract: The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution planes, and multi-agent collabora…

    arxiv.org3 days agoView details

  75. Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection

    arXiv:2609.13785v1 Announce Type: new Abstract: Oracle-style quantities, including virtual best solvers, selected-portfolio VBS, virtual-best encodings, and best-in-family summaries, are widely reported as upper bounds on what a deployable selector could achieve. In decomposed algorithm selection, an analogous partiti…

    arxiv.org3 days agoView details

  76. Do Not Restart: Residual Completion for Stateful Agent Handoffs

    arXiv:2609.13800v1 Announce Type: new Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinished obligations. We formulate this as commitment-constrained residual completion and introduce Commitment…

    arxiv.org3 days agoView details

  77. UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

    arXiv:2609.13849v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinfo…

    arxiv.org3 days agoView details

  78. ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

    arXiv:2609.13860v1 Announce Type: new Abstract: Querying clinical trial registries remains a manual and error-prone process, requiring researchers to navigate large volumes of semi-structured data without support for natural language interaction or cross-source synthesis. To address this, we introduce ClinAgent, a con…

    arxiv.org3 days agoView details

  79. SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery

    arXiv:2609.13945v1 Announce Type: new Abstract: Natural-language descriptions of optimization problems may be incomplete or vague about numerical information that a solver requires, including costs, capacities, demands, bounds, and penalties. A language model can translate the description into code, but when a require…

    arxiv.org3 days agoView details

  80. Synthetic Data in Marketing Research: How to Evaluate and When to Trust

    arXiv:2609.13995v1 Announce Type: new Abstract: Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic responde…

    arxiv.org3 days agoView details

  81. MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

    arXiv:2609.14399v1 Announce Type: new Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple com…

    arxiv.org3 days agoView details

  82. Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

    arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-depar…

    arxiv.org3 days agoView details

  83. Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement

    arXiv:2609.13466v1 Announce Type: new Abstract: Enterprise AI adoption has reached 78% of organizations globally, yet the infrastructure to govern that adoption has not kept pace. This paper identifies and characterizes the attestation deficit, a structural condition in which organizations maintain governance policies…

    arxiv.org3 days agoView details

  84. Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

    arXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by dynamically combining redundant features, s…

    arxiv.org3 days agoView details

  85. Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports

    arXiv:2609.14226v1 Announce Type: new Abstract: Automated radiology report generation has advanced rapidly in diagnostic accuracy, yet generated reports frequently diverge from the stylistic conventions of authentic radiologist writing in structure, diction, and uncertainty language, a gap which has direct implication…

    arxiv.org3 days agoView details

  86. Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation

    arXiv:2609.14015v1 Announce Type: new Abstract: Schizophrenia is a serious psychiatric disorder that affects millions of people worldwide, and its diagnosis remains primarily dependent on clinical assessment. Electroencephalography (EEG) provides a non-invasive approach to investigate brain activity and has shown pote…

    arxiv.org3 days agoView details

  87. Windowed A-K-MDP

    arXiv:2609.13676v1 Announce Type: new Abstract: Markov decision processes (MDPs) are used to support decision-making in conservation of biodiversity, but policies, even over small state spaces, can be difficult to interpret for conservation managers. K-MDP methods address this problem by building simpler MDPs with at…

    arxiv.org3 days agoView details

  88. GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents

    arXiv:2609.13667v1 Announce Type: new Abstract: Geospatial agents are increasingly expected to support recurring and evolving analytical tasks rather than execute isolated workflows. In such settings, effective agents must distill prior execution experience into reusable geospatial procedural knowledge to guide future…

    arxiv.org3 days agoView details

  89. Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

    arXiv:2609.13454v1 Announce Type: new Abstract: Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncer…

    arxiv.org3 days agoView details

  90. FedV-KGQA in Practice: Design Lessons and an Interactive Prototype

    arXiv:2609.13661v1 Announce Type: new Abstract: Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, facts are often held by organizations that share entity identifiers but own disjoint relation types, so no single party sees a complete reasoning chain. This poster…

    arxiv.org3 days agoView details

  91. From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements

    arXiv:2609.13535v1 Announce Type: new Abstract: The EU AI Act introduces mandatory requirements for high-risk AI systems with the explicit goal of ensuring the development and operation of trustworthy AI. At the same time, AI risk management practices rely on structured risk taxonomies to systematically identify and t…

    arxiv.org3 days agoView details

  92. PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems

    arXiv:2609.13152v1 Announce Type: new Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterati…

    arxiv.org3 days agoView details

  93. Learning to Refer from Estimated Listener Gaze

    arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are…

    arxiv.org3 days agoView details

  94. E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

    arXiv:2609.14302v1 Announce Type: new Abstract: Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through r…

    arxiv.org3 days agoView details

  95. Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

    arXiv:2609.13566v1 Announce Type: new Abstract: Industrial maintenance systems involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational cost using a single decision framework. While recent work has focused on reinforcement learning (RL) for maintenance sche…

    arxiv.org3 days agoView details

  96. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

    arXiv:2609.13463v1 Announce Type: new Abstract: The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders huma…

    arxiv.org3 days agoView details

  97. Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

    arXiv:2609.13980v1 Announce Type: new Abstract: Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, and frontier systems increasingly saturat…

    arxiv.org3 days agoView details

  98. LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents

    arXiv:2609.13437v1 Announce Type: new Abstract: Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, le…

    arxiv.org3 days agoView details

  99. Homeostatic Continual Learning

    arXiv:2609.13771v1 Announce Type: new Abstract: In this paper, I formulate a Continual Learning problem and propose a method named "Homeostatic Continual Learning" that enables an AI agent to learn continuously in a changing environment without catastrophic forgetting. The core of the method is to find outliers in the…

    arxiv.org3 days agoView details

  100. When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

    arXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with hu…

    arxiv.org3 days agoView details

  101. DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking

    arXiv:2609.13808v1 Announce Type: new Abstract: Structured knowledge fact checking aims to determine the truthfulness of natural language claims by reasoning over structured evidence. Recent program-generation approaches leverage large language models (LLMs) to generate executable graph reasoning programs, achieving s…

    arxiv.org3 days agoView details

  102. SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

    arXiv:2609.14320v1 Announce Type: new Abstract: Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the…

    arxiv.org3 days agoView details

  103. TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

    arXiv:2609.13158v1 Announce Type: new Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasonin…

    arxiv.org3 days agoView details

  104. Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability

    arXiv:2609.13869v1 Announce Type: new Abstract: Automatic sentence function identification is important for many downstream natural language processing (NLP) applications such as dialogue systems, text-to-speech synthesis, and machine translation. However, benchmark resources for Bangla sentence function classificatio…

    arxiv.org3 days agoView details

  105. Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself

    arXiv:2609.13657v1 Announce Type: new Abstract: Traditional recommender systems are typically trained to predict what item users will interact with next, but not why. However, offering personalized evidence for why a user might like the predicted item is an important way to enhance the service and to raise the likelih…

    arxiv.org3 days agoView details

  106. FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

    arXiv:2609.13580v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities across a wide range of natural language processing tasks. However, conventional fine-tuning typically relies on centralized data collection, bringing in privacy concerns. Federated learning (FL) enables c…

    arxiv.org3 days agoView details

  107. TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

    arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for question-answering tasks. However, these models often fail to capture dynamic temporal patterns, providing only implicit r…

    arxiv.org3 days agoView details

  108. ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

    arXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or ou…

    arxiv.org3 days agoView details

  109. When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting

    arXiv:2609.13709v1 Announce Type: new Abstract: Human corrections identify editable spans, but the examples receiving corrections may come from a selective feedback channel. We analyze this interaction at a fixed model checkpoint by decomposing a localized gradient into edited and retained untouched components. Square…

    arxiv.org3 days agoView details

  110. Formal Properties of Language as Constraints on Neural Dynamics

    arXiv:2609.14384v1 Announce Type: new Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural activity. Yet predictive success leaves mechanisms under-constrained. H…

    arxiv.org3 days agoView details

  111. Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation

    arXiv:2609.13581v1 Announce Type: new Abstract: Discharge summaries are lengthy medical documents that summarize a hospital in-patient visit. Automatically generating them can reduce documentation burden and return clinician time to patient care. Whereas Large Language Model (LLMs) could be used for this task, their A…

    arxiv.org3 days agoView details

  112. Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

    arXiv:2609.13445v1 Announce Type: new Abstract: Speech-to-speech LLMs like Moshi, and its derivative PersonaPlex, can listen and speak concurrently through full-duplex generation. However, they can begin speaking inappropriately during prolonged user silence: under digital-zero input, Moshi and PersonaPlex initiate sp…

    arxiv.org3 days agoView details

  113. Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories

    arXiv:2609.13154v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et al., 2022) and in-context learning (Brown et al., 2020) frequently push real-world prompts past several thousand tokens…

    arxiv.org3 days agoView details

  114. ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

    arXiv:2609.13356v1 Announce Type: new Abstract: In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric cap…

    arxiv.org3 days agoView details

  115. VeriDx: Earning the Right to Diagnose with Disease-Centric Verification

    arXiv:2609.14018v1 Announce Type: new Abstract: A correct diagnosis can still be reached for the wrong reasons. In clinical reasoning, every disease hypothesis creates obligations: key evidence must be checked, alternatives must be ruled out, contradictions must be resolved, useful tests must be considered, and closur…

    arxiv.org3 days agoView details

  116. LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems

    arXiv:2609.13805v1 Announce Type: new Abstract: In the era of the Internet of Things (IoT), coordinating connected electric vehicle (EV) charging scheduling to balance EV charging satisfaction, station profitability, and smart grid stability presents a complex multi-objective challenge. Existing Multi-Agent Reinforcem…

    arxiv.org3 days agoView details

  117. Cost Characterization of Vertically Partitioned Federated Knowledge Graphs

    arXiv:2609.13664v1 Announce Type: new Abstract: Knowledge graphs are increasingly distributed across autonomous organizations that share an entity space but own disjoint subsets of relations, forming a vertical partition. Answering a multi-hop query may require combining facts from several silos, making the partitioni…

    arxiv.org3 days agoView details

  118. Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation

    arXiv:2609.13519v1 Announce Type: new Abstract: Molecular design is most effective when generation mirrors the edits chemists actually make: extending a scaffold, replacing a substituent, or decorating a scaffold at a specified attachment site while optimizing molecular properties. Fragment-based molecular design natu…

    arxiv.org3 days agoView details

  119. Token Efficient Task Execution via Application Behavior Modeling for Web Agents

    arXiv:2609.13491v1 Announce Type: new Abstract: The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in…

    arxiv.org3 days agoView details

  120. Dynamic Learning Solutions: A System for Personalized Educational Video Generation

    arXiv:2609.14408v1 Announce Type: new Abstract: We present an automated pipeline that converts NCERT textbooks into interactive video explanations that respond directly to user queries. A user uploads a PDF and asks a question; the system then generates a video-based explanation as output, handling both text and visua…

    arxiv.org3 days agoView details

  121. Convergent Emergence of In-Context Learning Across Modalities

    arXiv:2609.14011v1 Announce Type: new Abstract: Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provided in its prompt and apply them to new inputs, has been extensively studied in large language models trained for next-token prediction on human text. R…

    arxiv.org3 days agoView details

  122. MANAS-2: Constrained Reconstruction for EEG Foundation Models

    arXiv:2609.13717v2 Announce Type: new Abstract: Masked reconstruction is widely used for EEG foundation models, but optimizing reconstruction on low-SNR waveforms does not necessarily produce the most useful latent representation. We introduce MANAS-2, a new EEG foundation model that combines a Raw-Band Hybrid (RBH) m…

    arxiv.org3 days agoView details

  123. Editorial routing shapes how computational results are qualified in AI-assisted scientific writing

    arXiv:2609.14288v1 Announce Type: new Abstract: Large language models increasingly analyze computational results and draft manuscripts, making reliable communication as important as correct analysis. Using fixed computational evidence, we tested whether assigning comparisons across modeling choices elsewhere in a rese…

    arxiv.org3 days agoView details

  124. How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition

    arXiv:2609.13747v1 Announce Type: new Abstract: Large language models solve hard problems through intermediate computations across multi-step reasoning. Traditional chain-of-thought encodes these computations as tokens. Recent continuous and recurrent methods instead move partial computations into fixed-dimensional la…

    arxiv.org3 days agoView details

  125. Converge Then Diversify: Decoupling Convergence and Diversity in Multi-Objective Bayesian Optimisation

    arXiv:2609.13396v1 Announce Type: new Abstract: Multi-objective Bayesian optimisation (MOBO) is a sample-efficient approach for optimising expensive black-box functions with multiple objectives. In MOBO, the goal is to adequately approximate the Pareto front; that is, to obtain a high-quality solution set with 1) good…

    arxiv.org3 days agoView details

  126. Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

    arXiv:2609.13406v1 Announce Type: new Abstract: When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances e…

    arxiv.org3 days agoView details

  127. PolicyMem: Geometric Policy Memory for LLM Governance

    arXiv:2609.13734v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy b…

    arxiv.org3 days agoView details

  128. The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation

    arXiv:2609.13615v1 Announce Type: new Abstract: For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bislama, and Solomon Pijin, with particular attention to broad domain performance. After pre-training on a large collectio…

    arxiv.org3 days agoView details

  129. AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

    arXiv:2609.13548v1 Announce Type: new Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for construc…

    arxiv.org3 days agoView details

  130. Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

    arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-i…

    arxiv.org3 days agoView details

  131. OrchSLM: Probing the Dynamics of Small Language Model Orchestration

    arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. S…

    arxiv.org3 days agoView details

  132. Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks

    arXiv:2609.13712v1 Announce Type: new Abstract: Many real-world scenarios can be represented using graph-structured data. However, traditional GNNs that transmit messages based on first-order neighbors have long faced several fundamental contradictions: increasing depth leads to over-smoothing, long-range dependencies…

    arxiv.org3 days agoView details

  133. Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation

    arXiv:2609.13238v1 Announce Type: new Abstract: Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during de…

    arxiv.org3 days agoView details

  134. RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines

    arXiv:2609.13389v1 Announce Type: new Abstract: Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specificat…

    arxiv.org3 days agoView details

  135. CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

    arXiv:2609.13413v1 Announce Type: new Abstract: We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 l…

    arxiv.org3 days agoView details

  136. A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics

    arXiv:2609.13561v1 Announce Type: new Abstract: Efficient utilization of supply chain analytics for decision making remains a significant challenge for planners, as critical tasks such as database querying, key performance indicator (KPI) analysis, demand forecasting, and performance diagnosis require heterogeneous ex…

    arxiv.org3 days agoView details

  137. Causal multi-modal AI for personalized chemosensitivity prediction

    arXiv:2609.13567v1 Announce Type: new Abstract: Chemotherapy improves survival for some patients with breast cancer, but doctors cannot reliably predict who. Current guidelines rely on recurrence scores as a proxy for treatment benefit, which may contribute to the overprescription of chemotherapy. Here we present a ca…

    arxiv.org3 days agoView details

  138. How User-AI Mistreatment Occurs and Matters in Conversational Systems?

    arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world dep…

    arxiv.org3 days agoView details

  139. A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text

    arXiv:2609.13481v1 Announce Type: new Abstract: Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extr…

    arxiv.org3 days agoView details

  1. Agnes-AI/Agnes-3.0-Flash

    image-text-to-text · transformers · safetensors · agnes

    huggingface.co4 days ago207 ptsView details

  2. oruk/orukeet

    automatic-speech-recognition · nemo · onnx · gguf

    huggingface.co4 days ago36 ptsView details

  3. coolthor/MiniMax-H3-pruned-NVFP4

    image-text-to-video · diffusion-single-file · text-to-video · image-to-video

    huggingface.co4 days ago22 ptsView details

  4. mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit

    text-generation · mlx · safetensors · qwen3_5_moe

    huggingface.co4 days ago17 ptsView details

  5. Chaman1234/Sparse-AST-BWM

    text-generation · pytorch · safetensors · sparse-ast

    huggingface.co4 days agoView details

  6. takanori-ishikawa/Qwen3.5-ANE-CoreML

    text-generation · coreml · ane · apple-neural-engine

    huggingface.co4 days agoView details

  1. openai/openai-python v3.14.0

    ## [3.14.0](https://github.com/openai/openai-python/compare/v3.13.0...v3.14.0) (2026-09-14) ### Features * **streaming:** normalize errors raised while reading streams ([#3827](https://github.com/openai/openai-python/issues/3827)) ([d7c41ef](https://github.com/openai/openai-pyth…

    github.com3 days agoView details

  2. ggml-org/llama.cpp b10970

    <details open> HIP: fattn-mma: use fp32 accumulation on MFMA devices (#28576) use fp32 accumulators in fattn-mma on CDNA </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47450174> **macOS/iOS:** - [macOS Apple…

    github.com3 days agoView details

  3. ggml-org/llama.cpp v0.4.1

    ## Overview llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0. ### API changes - Changed `llama_sampler_chain_n()` to return `int32_t` ins…

    github.com3 days agoView details

  4. ggml-org/llama.cpp b10952

    <details open> sycl : fix oneDNN scratchpad breaking the pool free order (#28704) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47287268> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggm…

    github.com4 days agoView details