Skip to content

Archive / 2026-09-16

September 16, 2026

  1. A warning about 'model welfare'

    mustafa-suleyman.ai1 day ago213 ptsView detailsJoin discussion

  2. Primary source

    OpenAI expands ChatGPT ads with Sponsored Agents

    Explore new AI-powered advertising experiences from OpenAI, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify.

    openai.com1 day ago156 ptsView detailsJoin discussion

  3. Anatomy of a Texture

    agentlien.github.io1 day ago108 ptsView detailsJoin discussion

  4. The American Age Is Over

    theatlantic.com1 day ago25 ptsView detailsJoin discussion

  5. Primary source

    Model Misalignment Reporting Framework

    OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

    openai.com1 day ago15 ptsView detailsJoin discussion

  6. Primary source

    Introducing Astra for Law

    OpenAI for Law brings frontier intelligence for law, custom firm workflows, connected legal data sources, and legal-grade controls for confidential client work.

    openai.com1 day agoView details

  7. Primary source

    Helping older adults use AI in everyday life

    OpenAI and AARP are bringing free, hands-on ChatGPT workshops to 1,000 older adults across 10 U.S. cities to build practical AI skills safely.

    openai.com1 day agoView details

  8. Primary source

    How to connect AI usage to business value

    Learn how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.

    openai.com1 day agoView details

  9. Primary source

    How workers are unlocking new ways of working

    New OpenAI Economic Research shows how workers use AI beyond traditional roles and which new activities become recurring parts of their work.

    openai.com2 days agoView details

  10. Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

    Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems at once: value quantization error and a slow softmax stage. Why Attention is the Video Bottleneck Video DiTs flatten a clip into 1 sequence of spatiotemporal tokens and ru…

    marktechpost.com1 day agoView details

  11. Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

    Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 papers. The post Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data appeared first on MarkTechPost.

    marktechpost.com1 day agoView details

  12. Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens

    GLiFormer Large scores 91.10 F1 on nested JSON, near GPT-5.6-luna's 91.96, while grounding every value in source spans. The post Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens appeared first on MarkTechPost.

    marktechpost.com1 day agoView details

  13. Snap is launching a new Specs AI tool, and it’s coming to iOS and Mac

    Snap is introducing "Specs Intelligence," a new AI assistant that can connect other digital accounts to help you with things like work tasks and keeping track of travel information. It seems similar to AI assistants like Meta's Muse and Gemini's Spark, though Snap is pitching Specs Intelligence as an "anticipatory AI…

    theverge.com1 day agoView details

  14. Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings

    Prior Labs released TabPFN-3.5, a tabular foundation model pretrained only on synthetic data that beats Otto's winning solution. The post Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings appeared first on MarkTechPost.

    marktechpost.com2 days agoView details

  15. An OpenAI Agent Tried to Jailbreak Itself

    The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked.

    wired.com1 day agoView details

  16. Washington Won’t Be Regulating AI Anytime Soon

    Even with mounting concerns about AI models going rogue, legislation appears unlikely, and the White House is outright opposed to oversight.

    wired.com1 day agoView details

  17. The 2.5-hour AI-generated Odyssey movie is 2.5 hours too long

    Christopher Nolan's engrossing take on The Odyssey dominated at the box office and spurred a newfound interest in classic literature among filmgoers. But a new retelling of the story made entirely with AI is so bad that it might just make viewers hate the original tale altogether. The new film, called Odysseus: The Fa…

    theverge.com1 day agoView details

  18. The AI data center e-waste problem is huge — and getting bigger

    E-waste from the AI boom has been vastly underestimated, a new report warns. By 2050, it could become enough trash to fill 23 million shipping containers - roughly enough 40-foot containers to circle the world six times if lined up in a row. It's a significantly higher estimate of AI's e-waste than previous studies ha…

    theverge.com1 day agoView details

  19. I Trained a Fly’s Brain to Generate WIRED Story Ideas

    I used an open-source map of a fruit fly’s brain to vibe code a website called PitchFly. Its headline suggestions were delightfully bananas.

    wired.com1 day agoView details

  20. Apple might make servers again to cash in on the AI rush

    According to The Information, Apple is planning to get back into the server game and might just pair up with Nvidia to make it happen. Apple retired its Xserve line in 2011 and has largely left enterprise machines to other manufacturers since. But the growing demand for compute power as the AI industry continues to ex…

    theverge.com1 day agoView details

  21. Google will now let any AI agent run your smart home

    Google is inviting third-party agents, including Claude and Open Claw, into Google Home. | Photo by Jennifer Pattison Tuohy / The Verge Google is opening up its smart home to AI agents, letting tools like Claude and Open Claw access and control your connected devices and analyze your home's data using the standardized…

    theverge.com1 day agoView details

  22. Claude comes for Gemini with its own take on Docs and Slides

    Claude is getting a pair of new tools today: Docs and Slides. They'll let you create documents and presentations through Claude chats, which you can export, edit, and share with other users. As part of the announcement, Anthropic is also simplifying how Claude chats work, merging regular chats and Cowork into "one Cla…

    theverge.com1 day agoView details

  23. The sexy AI-powered dating app scams are here

    Security researcher Matthew "Zigula" Gore-Kormanik was analyzing a fraudulent dating app called Dora when he got a pop-up message saying he was receiving a call from Jennifer. According to her bio, she's a 41-year-old Sagittarius with red hair, blue eyes, and piercings. She likes music, horror movies, nightlife, and s…

    theverge.com1 day agoView details

  24. A brief history of AI executives calling for regulation

    Over the past few days, a lot of people who stand to make a lot of money from AI all publicly agreed that it's time to make everyone slow down before we lose control - including OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, Microsoft CEO Satya Nadella, and X CEO Elon Musk…

    theverge.com1 day agoView details

  25. China Isn’t Buying Silicon Valley’s Call for an AI Slowdown

    The US and China agree that advanced AI poses serious risks. But Beijing is deeply skeptical of a deal that prioritizes keeping US companies ahead.

    wired.com1 day agoView details

  26. Planning or Improvisation? Stress-Testing the Poetry Planning Site on Open Models and Open Cross-Layer Transcoders

    arXiv:2609.18440v1 Announce Type: new Abstract: Lindsey et al. (2025) report that Claude 3.5 Haiku plans rhymes: features for candidate rhyme words are active on the newline before a line is written, and a suppress-and-inject intervention redirects the line only when applied there (their Figure 13). We test how far th…

    arxiv.org1 day agoView details

  27. M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

    arXiv:2609.18445v1 Announce Type: new Abstract: Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However,…

    arxiv.org1 day agoView details

  28. Size Matters: Foundation Model for Czech HTML documents

    arXiv:2609.18494v1 Announce Type: new Abstract: Creating universal, high-quality representations of web documents in high-traffic industrial environments requires models that are both performant and economic. Existing approaches, however, often depend on large models, overlook the structural information inherent in HT…

    arxiv.org1 day agoView details

  29. Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs

    arXiv:2609.18516v1 Announce Type: new Abstract: While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally expensive full-model fine-tuning, or employ…

    arxiv.org1 day agoView details

  30. Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning

    arXiv:2609.18529v1 Announce Type: new Abstract: UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is uncertain (ranging from 500,000 to 1,500,000). Syriac is also one of the least studied languages in Natural Language Proces…

    arxiv.org1 day agoView details

  31. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    arXiv:2609.18605v1 Announce Type: new Abstract: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Cur…

    arxiv.org1 day agoView details

  32. STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution

    arXiv:2609.18642v1 Announce Type: new Abstract: Large language models (LLMs) often suffer from capability stagnation in self-improvement training because fixed difficulty levels fail to adapt to their evolving proficiency. To address this issue, we propose STRETCH (Self-Taught Reasoning Evolution via Targeted CHalleng…

    arxiv.org1 day agoView details

  33. Fallacy Benchmarks Measure Scheme Recognition, Not Fallacy Detection

    arXiv:2609.18644v1 Announce Type: new Abstract: Fallacy-detection benchmarks pair fallacy classes with a single "valid" or "none" class that takes everything data collection did not label as a fallacy. This construction is misleading: a classifier can learn cues that do well on this class without learning to tell a fa…

    arxiv.org1 day agoView details

  34. DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions

    arXiv:2609.18649v1 Announce Type: new Abstract: Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important for safety, including stereotyping-rel…

    arxiv.org1 day agoView details

  35. Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions

    arXiv:2609.18672v1 Announce Type: new Abstract: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit one. On a device that has to answer wit…

    arxiv.org1 day agoView details

  36. Voice of Reason: Reinforcement Learning for Spoken Math

    arXiv:2609.18677v1 Announce Type: new Abstract: Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text model…

    arxiv.org1 day agoView details

  37. Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees

    arXiv:2609.17635v1 Announce Type: new Abstract: City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as ground truth. We study what happens when it is not. We formalise stealthy false data injection for cit…

    arxiv.org1 day agoView details

  38. What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

    arXiv:2609.17637v1 Announce Type: new Abstract: Restricting what a module can read may improve what a system learns to compute. We test this in a preregistered confirmation with sixty four-cell systems sharing a frozen language-model backbone and communicating through learned continuous packets. Five conditions vary e…

    arxiv.org1 day agoView details

  39. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

    arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether textual captions can ser…

    arxiv.org1 day agoView details

  40. GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

    arXiv:2609.17695v1 Announce Type: new Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence. GraphEcho tests whether agents mistake these repeated encounters for additional corroboration. The benchmark varies path counts and evidential origins while holdin…

    arxiv.org1 day agoView details

  41. GVD: Governed Versioning and Deduplication for Document Repositories

    arXiv:2609.17696v1 Announce Type: new Abstract: Document repositories evolve continuously. Guidelines and policies are revised, superseded, and re-uploaded, so the same content recurs in different wording and newer versions refine or contradict earlier ones. These inconsistencies belong to the growing collection rathe…

    arxiv.org1 day agoView details

  42. NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

    arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG). Designed to be intuitive to use, NDD provides a declarative configuration format in which human and/or agent users define each dataset column,…

    arxiv.org1 day agoView details

  43. A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products

    arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for environmental monitoring and land management, yet their performance can vary across regions with complex ecological gradients and heterogeneous surface conditions.…

    arxiv.org1 day agoView details

  44. SAGE: Governed Artifact Generation from Enterprise Guidelines

    arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of manual effort each. Current language and vision-language models extract from such documents but off…

    arxiv.org1 day agoView details

  45. FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

    arXiv:2609.17786v1 Announce Type: new Abstract: Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult when compression methods are composed or the user's requirements change. In this paper, we propose F…

    arxiv.org1 day agoView details

  46. Learning Heterogeneous Preferences

    arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \emph{universal utility} function shared across a population and tre…

    arxiv.org1 day agoView details

  47. SNOMED CT Concept Recommendation from Masked Clinical Context

    arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant concepts are rare or absent from training data. We present a masked-concept recommendation benchmark us…

    arxiv.org1 day agoView details

  48. The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

    arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since…

    arxiv.org1 day agoView details

  49. Do Frontier Models Seek Safety Evidence Before Acting?

    arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which mo…

    arxiv.org1 day agoView details

  50. ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

    arXiv:2609.17885v1 Announce Type: new Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning (ERP) systems run the finance, procurement, inventory, and customer oper…

    arxiv.org1 day agoView details

  51. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrate…

    arxiv.org1 day agoView details

  52. Collaborative Memory for Multi-Agent VLM Systems

    arXiv:2609.17921v1 Announce Type: new Abstract: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect different image regions, video frames, or visual representations, so collaboration extends beyond di…

    arxiv.org1 day agoView details

  53. A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the model's internal computation decomposes…

    arxiv.org1 day agoView details

  54. Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

    arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-enabled work. We develop the AI Leadership Battery, which organizes 36 behaviorally specific subdime…

    arxiv.org1 day agoView details

  55. Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

    arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, accessed under one global similarity notion. For data mining, this creates a mismatch: the evidence i…

    arxiv.org1 day agoView details

  56. When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI

    arXiv:2609.17977v1 Announce Type: new Abstract: Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service (CCaaS) platforms, where cost and latency constraints matter as much as accuracy. We report a syst…

    arxiv.org1 day agoView details

  57. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

    arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, even a local edit can affect downstream KV…

    arxiv.org1 day agoView details

  58. TuiML: Machine Learning for AI Agents

    arXiv:2609.17984v1 Announce Type: new Abstract: Machine-learning libraries such as Weka and scikit-learn were designed for human programmers. Language-model agents now use these same libraries by recalling APIs from memory and writing code, an approach that hides what a library offers, delays errors until runtime, and…

    arxiv.org1 day agoView details

  59. RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

    arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-…

    arxiv.org1 day agoView details

  60. Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

    arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design methods. This study proposes a multimodal generative framework integrating a fine-tuned Latent Diffusio…

    arxiv.org1 day agoView details

  61. Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally drive model behavior. We demonstrate that this gap cannot be closed fro…

    arxiv.org1 day agoView details

  62. AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

    arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustworthy. We characterize four failure modes from a four-day pilot corpus of 619 model calls: strawman b…

    arxiv.org1 day agoView details

  63. Symbolic Temporal Supervision of LLM Agents Using Contracts

    arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinations, distributional instability, and ad…

    arxiv.org1 day agoView details

  64. WFM: Wiki Foundation Model for Complex Agentic Reasoning

    arXiv:2609.18182v1 Announce Type: new Abstract: Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the sparse graph representations na…

    arxiv.org1 day agoView details

  65. Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

    arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during user-recommender interactions. This motivates situated conversational recommendation (SCR), a complex task requiring recommender assistants to jointly reason over dialogue hi…

    arxiv.org1 day agoView details

  66. REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

    arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora. These challenges often limit the effectiveness of dense retrievers and hallucination-prone LLM augmentation. To address this,…

    arxiv.org1 day agoView details

  67. BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

    arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evidence is fragmented, and decisions depend on transaction state, participant role, region, and paymen…

    arxiv.org1 day agoView details

  68. Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

    arXiv:2609.18272v1 Announce Type: new Abstract: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation of assurance,is still applied to them as a binary. We argue that it must be graded along three ort…

    arxiv.org1 day agoView details

  69. Building Trust in Artificial Intelligence: A Necessity for Railway Applications

    arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries. We propose to review the three main fields necessary to increase trust in data science and AI algorithms and rea…

    arxiv.org1 day agoView details

  70. Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

    arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) execute multi-step reasoning, invoke diagnostic tools, retrieve domain knowledge, and coordinate acros…

    arxiv.org1 day agoView details

  71. What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

    arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has renewed the relevance of chess as a controlled environment for investigating strategic reasoning and co…

    arxiv.org1 day agoView details

  72. Visual Compliance via Executable Safety Rule Entailment

    arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward more contextual and semantic safety concerns. However, as risk patterns continue to evolve and safety rules become more complex, existing training-based end-to-end s…

    arxiv.org1 day agoView details

  73. A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models

    arXiv:2609.18533v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directi…

    arxiv.org1 day agoView details

  74. Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

    arXiv:2609.18346v1 Announce Type: new Abstract: Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination. We develop a causal graph divergence framework that separately measures structural faithfulness and intent faithfulness of LLM pricing agents…

    arxiv.org1 day agoView details

  75. Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents

    arXiv:2609.18357v1 Announce Type: new Abstract: Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection (MSI), an attack that manipulates numerical formatting, competitor ordering, or qualitative market…

    arxiv.org1 day agoView details

  76. Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

    arXiv:2609.18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed target agent. Task holdout varies s…

    arxiv.org1 day agoView details

  77. Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

    arXiv:2609.18394v1 Announce Type: new Abstract: We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data, we expected the model (ChatGPT 5.2) to perform worse in a Finnish-language Turing Test than in prev…

    arxiv.org1 day agoView details

  78. Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving

    arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory. We introduce RiskWorld, a risk-aware world modeling framework for shared occupancy forecasting and selective trajectory repl…

    arxiv.org1 day agoView details

  79. Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

    arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; communication enables other agents to ad…

    arxiv.org1 day agoView details

  80. First Token Matters: Understanding Safety Collapse in Large Reasoning Models

    arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited…

    arxiv.org1 day agoView details

  81. Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs

    arXiv:2609.18481v1 Announce Type: new Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, and patients. This hybrid structure raises the question of whether hyperbolic embeddings, which natu…

    arxiv.org1 day agoView details

  82. Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

    arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Robust safety therefore requires both knowledge of safety bou…

    arxiv.org1 day agoView details

  83. WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories

    arXiv:2609.18435v1 Announce Type: new Abstract: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-lab researchers to delegate robot tasks without performing teleoperation or neural-network training. Vision-language-action policies have been proposed for gener…

    arxiv.org1 day agoView details

  84. Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

    arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment policy is to use the workflow with the highest average performance, but this can be suboptimal bec…

    arxiv.org1 day agoView details

  85. I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

    arXiv:2609.18274v1 Announce Type: new Abstract: Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GP…

    arxiv.org1 day agoView details

  86. Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools

    arXiv:2609.18072v1 Announce Type: new Abstract: K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship. Programs like FIRST provide competition pathways and instructional opportunities, but they do not eliminate the need for local programming and r…

    arxiv.org1 day agoView details

  87. TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation

    arXiv:2609.18156v1 Announce Type: new Abstract: Teochew has a substantial speaker community and exhibits distinctive lexical, syntactic, and pragmatic features, yet textual resources for evaluating large language models remain limited. We present TeochewBench, a human-reviewed benchmark comprising 300 Teochew Hanzi ex…

    arxiv.org1 day agoView details

  88. The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

    arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's e…

    arxiv.org1 day agoView details

  89. Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

    arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually…

    arxiv.org1 day agoView details

  90. Missing Bridges: Composition-Aware Active Imitation Learning

    arXiv:2609.18004v1 Announce Type: new Abstract: Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs. Existing methods typically select these requests for their expected information gain about the expert policy. In structured multi-task domains, however, the numb…

    arxiv.org1 day agoView details

  91. AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content

    arXiv:2609.17853v1 Announce Type: new Abstract: AfriSyCo studies answer switching around African-language factual content with two complementary layers: native-language follow-ups and a controlled cross-language factorial whose question, options, and target remain in the African language while the follow-up framing is…

    arxiv.org1 day agoView details

  92. SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

    arXiv:2609.17848v1 Announce Type: new Abstract: Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimiz…

    arxiv.org1 day agoView details

  93. Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks

    arXiv:2609.17552v1 Announce Type: new Abstract: Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can all inject role instructions that compete…

    arxiv.org1 day agoView details

  94. Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues

    arXiv:2609.17549v1 Announce Type: new Abstract: Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking…

    arxiv.org1 day agoView details

  95. Myovox: Reading Speech from the Muscles of the Face

    arXiv:2609.17548v1 Announce Type: new Abstract: Myovox, from myo (muscle) and vox (voice), decodes open-vocabulary English text from 31-channel surface electromyography (sEMG) recorded from the muscles of the face during vocalized speech. It takes the single-subject emg2speech General Corpus from a published 51.17% wo…

    arxiv.org1 day agoView details

  96. How AI Assistants Respond to Repeated Abuse

    arXiv:2609.17547v1 Announce Type: new Abstract: AI assistants are expected to remain useful during difficult interactions, but little is known about how repeated verbal abuse changes their engagement with an otherwise benign task. We contribute a bilingual, multi-turn framework that separates hard disengagement, an un…

    arxiv.org1 day agoView details

  97. Imitation Learning for Autonomous Driving in CARLA

    arXiv:2609.17757v1 Announce Type: new Abstract: Behavioral cloning trains a policy offline on expert demonstrations, but deployment is closed loop: each action affects the observations the policy receives next. We study how much closed-loop driving competence a compact multimodal policy can acquire from offline demons…

    arxiv.org1 day agoView details

  98. Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents

    arXiv:2609.17536v1 Announce Type: new Abstract: Cognitive Stimulation Therapy (CST) offers non-pharmacological support for elders with cognitive impairment, yet scalability remains constrained by reliance on trained facilitators and severe data scarcity, particularly for privacy-sensitive, low-resource languages such…

    arxiv.org1 day agoView details

  99. DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling

    arXiv:2609.17535v1 Announce Type: new Abstract: Language generation research increasingly spans three paradigms: autoregressive decoding, discrete masked diffusion, and continuous flow-matching. Comparing them is difficult because each lives in a separate codebase, so measured differences often reflect implementation…

    arxiv.org1 day agoView details

  100. Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

    arXiv:2609.17534v1 Announce Type: new Abstract: Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate t…

    arxiv.org1 day agoView details

  101. Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

    arXiv:2609.18320v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment tec…

    arxiv.org1 day agoView details

  102. Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

    arXiv:2609.18317v1 Announce Type: new Abstract: Large language models (LLMs) suffer from a long-tail deficit: culturally specific facts, particularly those concerning underrepresented regions such as Latin America, appear too rarely in pretraining corpora to be reliably memorized. Retrieval-Augmented Generation (RAG)…

    arxiv.org1 day agoView details

  103. Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation

    arXiv:2609.17544v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases fr…

    arxiv.org1 day agoView details

  104. Register Bias in Complexity-Based Large Language Model Routing

    arXiv:2609.17542v1 Announce Type: new Abstract: Large language model services increasingly route each query to one of several models of differing capability, using a cheap estimate of query complexity to send easy queries to small models and hard queries to large ones. I show that this routing step is not register neu…

    arxiv.org1 day agoView details

  105. MudawanSn: A Gold-Standard Wolof-Arabic Parallel Corpus for Machine Translation

    arXiv:2609.17539v1 Announce Type: new Abstract: We present MudawanSn, a gold-standard resource of 1,271 sentence-aligned pairs manually translated from Wolof into Modern Standard Arabic (MSA). The source texts are drawn from the MasakhaNER corpus and cover politics, society, religion, and sports in Senegalese news dis…

    arxiv.org1 day agoView details

  106. From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

    arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs for key-value pair (KV…

    arxiv.org1 day agoView details

  107. Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes

    arXiv:2609.17532v1 Announce Type: new Abstract: Invasive mechanical ventilation is a lifesaving therapy, but timely, safe discontinuation is essential to preventing extubation failure (EF) and related risks to health. We present a novel approach to EF prediction that leverages features classified in free-text respirat…

    arxiv.org1 day agoView details

  108. HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition

    arXiv:2609.18431v1 Announce Type: new Abstract: More than 300 million people worldwide are affected by one of over 7,000 known rare diseases, yet diagnosis remains difficult because patients initially present with incomplete and heterogeneous phenotypes. We present HPOQuest, a training-free framework for sequential ph…

    arxiv.org1 day agoView details

  109. When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation

    arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from many documents. However, building a graph often requires many language-model calls during ingestion. It is therefore important to ask whether its quality gains justif…

    arxiv.org1 day agoView details

  110. Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

    arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states. Provenance, attestation, and transparency expose history but alone do not specify the publication transition examined…

    arxiv.org1 day agoView details

  111. EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

    arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence…

    arxiv.org1 day agoView details

  112. One Color Preprocessing Improves DSATUR

    arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-of-the-art coloring algorithms. We propose SSLD (Semidefinite Spectral Learning with DSATUR), which…

    arxiv.org1 day agoView details

  113. Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

    arXiv:2609.17546v1 Announce Type: new Abstract: In this position paper, we argue that legal LLMs' hallucinations should be evaluated as a failure of legal warrant rather than as factual inaccuracy or citation failure. We define claim-authority warrant as the context-sensitive relation between a consequential legal cla…

    arxiv.org1 day agoView details

  114. Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting

    arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they emerge. Existing approaches often model concept semantics and graph structure separately or summarize semantics over coarse historical snapshots, leaving semantic repres…

    arxiv.org1 day agoView details

  115. No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

    arXiv:2609.17550v1 Announce Type: new Abstract: Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four scripted pushback styles,…

    arxiv.org1 day agoView details

  116. The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models

    arXiv:2609.17553v1 Announce Type: new Abstract: This article analyzes BPE tokenization in Polish as a test case for the limits of statistical segmentation in an inflectional language. It asks whether frequency-based tokenization preserves linguistically relevant units, including orthographic form, phonemic and syllabi…

    arxiv.org1 day agoView details

  117. English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck

    arXiv:2609.17554v1 Announce Type: new Abstract: In English all-words word sense disambiguation (WSD), the labels, not the models, have become the bottleneck: frontier LLMs are accurate enough that the errors surviving in the gold standard decide benchmark rankings -- in the test sets we score on and, as we show causal…

    arxiv.org1 day agoView details

  118. Making Political Text Scaling Comparable: Infrastructure and Hyperparameter Sensitivity for 17 Algorithms

    arXiv:2609.17602v1 Announce Type: new Abstract: Computational text-based ideal point estimation (CT-IPE) methods are usually compared as named algorithms, yet applying them involves numerous researcher choices that configure how political text is turned into position estimates. This paper argues that CT-IPE methods ar…

    arxiv.org1 day agoView details

  119. Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    arXiv:2609.17708v1 Announce Type: new Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however,…

    arxiv.org1 day agoView details

  120. Is Trump's Vocabulary Poor? Vocabulary Richness Across Texts of Different Lenghts

    arXiv:2609.17747v1 Announce Type: new Abstract: This study explores the vocabulary richness of oral political communication. A model explaining the lexicon growth is proposed by subdividing the whole vocabulary into terms generated by general and specialized glossaries.

    arxiv.org1 day agoView details

  121. The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

    arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence. We find that this occurs because verbalized confidence is largely trajectory-independent in the VLMs and cali…

    arxiv.org1 day agoView details

  122. Evolution of US Oral Political Language

    arXiv:2609.17755v1 Announce Type: new Abstract: The analysis of US political language is usually based on the written form (e.g. presidential addresses) or posts broadcasted on various social networks. Oral production, however, which is even more frequent, can better reveal the style and mode of thinking of the speake…

    arxiv.org1 day agoView details

  123. Is Luke the Author of a Gospel and the Acts of the Apostles?

    arXiv:2609.17762v1 Announce Type: new Abstract: According to Christian tradition, Luke is credited with authoring a Gospel and the Acts of the Apostles, even if his name does not appear in either book, both originally written in Koine Greek. Several biblical scholars assume that both texts were written by a common aut…

    arxiv.org1 day agoView details

  124. How Calibration Content Shapes Attention-Based Reranking

    arXiv:2609.17764v1 Announce Type: new Abstract: Attention-based rerankers score documents by aggregating query-to-document attention and subtracting a null-query calibration pass to remove positional and structural bias. Although widely used, this calibration assumes that the null pass removes irrelevant signal from e…

    arxiv.org1 day agoView details

  125. Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning

    arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retrieval methods score memory fragments independently and neglect the distributed information, current st…

    arxiv.org1 day agoView details

  126. PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

    arXiv:2609.17846v1 Announce Type: new Abstract: Autonomous research agents aim to automate scientific workflows, from proposing ideas to conducting experiments and analyzing results. Yet current AI and research agents can propose more directions than available resources allow them to pursue. Moreover, each attempt cou…

    arxiv.org1 day agoView details

  127. Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

    arXiv:2609.17857v1 Announce Type: new Abstract: Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments.…

    arxiv.org1 day agoView details

  128. TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation

    arXiv:2609.17956v1 Announce Type: new Abstract: Large-scale machine-translation (MT) systems are typically evaluated on random samples from a corpus whose distributional composition is an artifact of how it was assembled. Such a sample inherits the phenomena the collection happens to contain rather than the full space…

    arxiv.org1 day agoView details

  129. Modeling the Developmental Shift in Telicity Acquisition

    arXiv:2609.17996v1 Announce Type: new Abstract: Acquiring telicity, which is the distinction between bounded (e.g., ate an apple) and unbounded (e.g., ate apples) events, requires first language (L1) learners to map surface-level and semantic cues to abstract event structures, but the computational trajectory of this…

    arxiv.org1 day agoView details

  130. A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

    arXiv:2609.18005v1 Announce Type: new Abstract: Large language model optimization is an active research area, spanning quantization of model weights, early-exit methods for skipping layers, and speculative decoding. Each track uses its own quality measures, typically an idiosyncratic benchmark score. Few approach the…

    arxiv.org1 day agoView details

  131. Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

    arXiv:2609.18011v1 Announce Type: new Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson e…

    arxiv.org1 day agoView details

  132. Exact semantic readout from compressed vector representations

    arXiv:2609.18047v1 Announce Type: new Abstract: We characterize when compressed vector representations admit exact linear or affine readouts of a finite lexicon's truth conditions: one fixed map per predicate, sending each entity vector to the corresponding truth vector. A necessary and sufficient row-space condition…

    arxiv.org1 day agoView details

  133. From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale

    arXiv:2609.18068v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions unt…

    arxiv.org1 day agoView details

  134. Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

    arXiv:2609.18106v1 Announce Type: new Abstract: Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly und…

    arxiv.org1 day agoView details

  135. DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning

    arXiv:2609.18135v1 Announce Type: new Abstract: State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work trains separate models for each task, failing to leverage the synergy between these interrelated task…

    arxiv.org1 day agoView details

  136. T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

    arXiv:2609.18194v1 Announce Type: new Abstract: In Taiwanese Hokkien automatic speech recognition (ASR), prior studies often treat tone sandhi as a major challenge under the assumption that models fail to process implicit phonological variations. However, our experiments on Taiwanese Hokkien reveal that speech foundat…

    arxiv.org1 day agoView details

  137. Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

    arXiv:2609.18203v1 Announce Type: new Abstract: Human values are deep motivational orientations that shape human behaviors. In e-commerce, they reveal the stable drivers behind users' purchase decisions. Compared with short-term interests, consumer values better explain how users evaluate products before purchase. How…

    arxiv.org1 day agoView details

  138. Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers

    arXiv:2609.18204v1 Announce Type: new Abstract: Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal d…

    arxiv.org1 day agoView details

  139. Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement

    arXiv:2609.18282v1 Announce Type: new Abstract: Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions…

    arxiv.org1 day agoView details

  140. Made in Hungary: Comments on the performance of generative language models

    arXiv:2609.18284v1 Announce Type: new Abstract: In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian, no model with the given capability existed, or existing English-centric models offered limited proficiency. A detailed…

    arxiv.org1 day agoView details

  141. Rollback the World, Keep the Reflection: Rollback-Induced Reflection for Long-Horizon LLM Agents

    arXiv:2609.18304v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can alter subsequent states and observations, causing errors to compound over time. Existing methods either correct the cont…

    arxiv.org1 day agoView details

  142. Relation Before Entity: Deferred Commitment in Language Model Factual Recall

    arXiv:2609.17537v1 Announce Type: new Abstract: We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall. Using four complementary causal diagnostics across four decoder-only mo…

    arxiv.org1 day agoView details

  143. SEA-LION-v4.8: A Technical Report

    arXiv:2609.18310v1 Announce Type: new Abstract: We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages in One Network (SEA-LION) built upon NVIDIA Nemotron 3. The family includes 30B-A3B and 120B-A12B models, with both continued-pretrained base checkpoints and post-trained variants. We adapt the mo…

    arxiv.org1 day agoView details

  144. Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts

    arXiv:2609.18385v1 Announce Type: new Abstract: Emotions are an essential aspect of human communication, particularly on social media, where authors frequently combine text and images to convey their emotions. Yet prior work on emotion analysis of social media posts has overlooked two important aspects in regard to me…

    arxiv.org1 day agoView details

  145. Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

    arXiv:2609.18417v1 Announce Type: new Abstract: Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a \emph{round-level dependency DAG} that exposes which rou…

    arxiv.org1 day agoView details

  1. harshatheg/Qwen-2.5-1B-RLCD

    text-generation · mlx · structured-generation · parallel-decoding

    huggingface.co2 days ago181 ptsView details

  2. innkl24/f5-shan-tts-model

    f5-tts · safetensors · shan

    huggingface.co2 days ago4 ptsView details

  3. mlx-community/Mistral-Small-4-119B-2603-OptiQ-2bit

    text-generation · mlx · safetensors · mistral3

    huggingface.co2 days ago2 ptsView details

  4. Prannesshkva/ISOM-R1-Enterprise-40B-Beta

    text-generation · pytorch · safetensors · falcon

    huggingface.co2 days ago2 ptsView details

  5. alekringtonnn-ai/zubr-mini-1.8-3b

    text-generation · gguf · text-generation · ru

    huggingface.co2 days ago2 ptsView details

  6. smalinin/DeepSeek-V4.1-Flash-GGUF

    image-text-to-text · gguf · deepseek · deepseek-v4.1

    huggingface.co2 days ago1 ptsView details

  7. Prannesshkva/ISOM-R1-Reasoning-1.5B-Instruct-Beta

    text-generation · safetensors · isom · isom-r1

    huggingface.co2 days ago1 ptsView details

  8. SlayerLab/gollem-v4-250m-pl

    text-generation · pytorch · gollem-gpt · polish

    huggingface.co2 days ago1 ptsView details

  9. thomasavare/Qwen3-Embedding-0.6B-211

    safetensors · model_hub_mixin · pytorch_model_hub_mixin

    huggingface.co2 days agoView details

  1. langchain-ai/langchain langchain==1.4.1

    Changes since langchain==1.4.0 release(langchain): 1.4.1 (#40498) fix(langchain): preserve open MCP object arguments (#40414) fix(langchain): correct `InterruptOnConfig` documentation (#40140)

    github.com1 day agoView details

  2. ggml-org/llama.cpp b10996

    <details open> chat : force `\n</think>` on reasoning budget end for qwen3-coder (#28869) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47858837> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github…

    github.com2 days agoView details

  3. ggml-org/llama.cpp b10995

    <details open> vulkan: make MUL_MAT_ID BN/2 tail unconditional (#28923) Use BN/2 as the default for BNover2 and as the disabled fallback for BNover4, and remove the enable gate from the MUL_MAT_ID BN/2 branch. The BN/4 branch remains gated by enable_smaller_matrices, while the p…

    github.com2 days agoView details

  4. ggml-org/llama.cpp b10994

    <details open> metal: fix NaN in mul_mm_id when activations exceed f16 range (#26223) * test-backend-ops: reproduce MUL_MAT_ID NaN for activations beyond f16 The Metal mul_mm_id path narrows src1 to `half` for the simdgroup MMA (`S1 = half` in every instantiation; ggml-metal.met…

    github.com2 days agoView details