Archive / 2026-09-17
September 17, 2026
News
View all news →A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
hacktron.ai6 hours ago334 ptsView detailsJoin discussion
Bend – A language that blocks AI mistakes via proof, on CPU and GPU
bend-lang.com12 hours ago442 ptsView detailsJoin discussion
github.com5 hours ago94 ptsView detailsJoin discussion
waymo.com5 hours ago107 ptsView detailsJoin discussion
Hister: A private search engine for the pages you visit and the files you keep
github.com16 hours ago589 ptsView detailsJoin discussion
fex-emu.com5 hours ago96 ptsView detailsJoin discussion
qwen.ai10 hours ago194 ptsView detailsJoin discussion
iankduncan.com12 hours ago205 ptsView detailsJoin discussion
More than 100k people in Japan are now aged 100 or older
bbc.com12 hours ago185 ptsView detailsJoin discussion
sockpuppet.org11 hours ago138 ptsView detailsJoin discussion
henriquenunez.eu11 hours ago137 ptsView detailsJoin discussion
How GLM built its own inference infrastructure
z.ai1 day ago393 ptsView detailsJoin discussion
martinfowler.com19 hours ago214 ptsView detailsJoin discussion
The American Religion of Self-Storage Facilities
newyorker.com20 hours ago226 ptsView detailsJoin discussion
How Uber Protects Against Retry Storms
uber.com12 hours ago91 ptsView detailsJoin discussion
Show HN: Share your AI Setup, Learn from others
mysetup.ai20 hours ago211 ptsView detailsJoin discussion
AI safety is mostly a sex cult
skywriter.blue1 day ago292 ptsView detailsJoin discussion
Rate limits on GitLab.com are changing
about.gitlab.com17 hours ago165 ptsView detailsJoin discussion
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arxiv.org16 hours ago141 ptsView detailsJoin discussion
LLM Classification Is Feature Engineering
minimallysufficient.com17 hours ago100 ptsView detailsJoin discussion
devin.ai6 hours ago20 ptsView detailsJoin discussion
Show HN: I built a new version of my fun spatial 3D online meeting app
flat.social20 hours ago94 ptsView detailsJoin discussion
institute.deepmind.com16 hours ago62 ptsView detailsJoin discussion
Microsoft, OpenAI lose fight to hide internal docs admitting scraping is theft
arstechnica.com13 hours ago44 ptsView detailsJoin discussion
How SpaceX Streamlined the Raptor Engine
construction-physics.com12 hours ago38 ptsView detailsJoin discussion
docs.aws.amazon.com7 hours ago15 ptsView detailsJoin discussion
The FAA's plan to fix air traffic? $875M worth of AI
techcrunch.com9 hours ago21 ptsView detailsJoin discussion
Microsoft exec called AI scraping 'the largest theft of labor in human
techcrunch.com9 hours ago20 ptsView detailsJoin discussion
Unix Was Right. We Made Linux Too Complicated [video]
youtube.com7 hours ago13 ptsView detailsJoin discussion
Canto: A speech model built for the real world
wisprflow.ai14 hours ago36 ptsView detailsJoin discussion
Hackers Used Anthropic's Claude to Break into OpenAI
wsj.com8 hours ago13 ptsView detailsJoin discussion
OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
asiaai.fyi17 hours ago40 ptsView detailsJoin discussion
Jev means structured output is interesting again
seangoedecke.com8 hours ago11 ptsView detailsJoin discussion
How, Exactly, Could A.I. Kill Us?
newyorker.com20 hours ago37 ptsView detailsJoin discussion
Eric S. Raymond – This is the case against AI Doom – pass it on
twitter.com11 hours ago14 ptsView detailsJoin discussion
Run QWEN3.8 27B on 16gb Nvidia GPUs
github.com13 hours ago17 ptsView detailsJoin discussion
Show HN: Craigslist for agent skills, curated by a human
skillbay.sh16 hours ago23 ptsView detailsJoin discussion
A Netflix documentary led to resurrecting a 9/11 memorial
gwintrob.com12 hours ago14 ptsView detailsJoin discussion
Priest, Monk, and Mathematician
logangraves.com16 hours ago20 ptsView detailsJoin discussion
pogueman.substack.com15 hours ago16 ptsView detailsJoin discussion
Show HN: AutoBot – live voice control for long-running AI work
github.com16 hours ago18 ptsView detailsJoin discussion
Plugin4Shell – Zero Click RCE Vulnerability found in top four coding agents
air.security13 hours ago11 ptsView detailsJoin discussion
Figure AI - Helix 2.5 Robot: Zero-Shot Home Generalization
figure.ai13 hours ago11 ptsView detailsJoin discussion
Show HN: MCPJam - the first testing & evaluations platform for MCP servers
mcpjam.com13 hours ago10 ptsView detailsJoin discussion
Show HN: Die With Me – Claude and Codex rate limits as AIM away messages
diewithme.co16 hours ago11 ptsView detailsJoin discussion
AI model watermarking changes agent behavior
theregister.com17 hours ago11 ptsView detailsJoin discussion
PC ports of old console games are the new AI vibe coding battleground
pcgamer.com1 day ago11 ptsView detailsJoin discussion
- Primary source
The future of practice: Enabling teachers to create learning interactives with generative UI
Education Innovation
research.google12 hours agoView details
Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
Microsoft's AKS engineering team open-sourced TauGrid on August 28, 2026, packaging the tau CLI, Kueue queueing, KubeRay orchestration, GPU node health monitoring and observability into one Helm install. It is MIT licensed and deployable now on any Kubernetes 1.30+ cluster with GPU nodes, kubectl and Helm 3.0 or later…
marktechpost.com11 hours agoView details
Anthropic redesigned Projects in Claude Code. The old project was a folder: some files plus one chat. The new one is a single ongoing conversation where Claude acts as coordinator. You describe work, and Claude decides what becomes a thread. Each thread is a full Claude Code cloud session running on its own branch and…
marktechpost.com12 hours agoView details
- Primary source
How Cooley is accelerating IPO work with ChatGPT
Cooley built GO Public with ChatGPT Work to bring intelligence to the IPO process, helping lawyers surface issues earlier and focus judgment where it matters most.
openai.com21 hours agoView details
Here’s What the AI Apocalypse Could Look Like
This week on “Uncanny Valley,” we discuss three possible AI doomsday scenarios, AI safety, and the unexpected bipartisan alliance forming against AI.
wired.com11 hours agoView details
The AI ‘Slowdown’ Is an Antitrust Mess
By framing their efforts as a “slowdown” rather than an industry-wide push for better security standards, AI labs may have set themselves up for years of regulatory headaches.
wired.com13 hours agoView details
The AI Superintelligence Slowdown
Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a number of leading US AI companies are publicly suggesting it’s time to pump the…
theverge.com13 hours agoView details
Claude Code relaunches Projects to manage multiple AI agents in the cloud
The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, each project has "threads" running different tasks in parallel, with a "coordinator…
theverge.com14 hours agoView details
The AI Slowdown Debate Crashed Salesforce’s Party
The Dreamforce conference became an unlikely battleground for the CEOs of OpenAI, Anthropic, and Nvidia to debate whether AI development should slow down.
wired.com14 hours agoView details
OpenAI can disclose misalignment before fixes exist. Its 6 initial reports include fabricated data and leaked API keys. The post OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training appeared first on MarkTechPost.
marktechpost.com1 day agoView details
Google Research has introduced Retrieve-for-Train (R4T), a framework for search that returns coherent, diverse result sets. It trains a fan-out language model with RL once, using groundedness, diversity, and alignment rewards. That model then synthesizes training data for a 53.9M-parameter diffusion retriever. The ret…
marktechpost.com1 day agoView details
AI is feared globally as the destroyer of jobs
In 34 of the 37 surveyed countries, people are more likely to believe AI will lead to job losses over the next 20 years. | Image: Pew Pew Research has published a new global survey that sheds light on how people view AI, including its impact on jobs, life in general, and income inequality. The survey questioned 42,151…
theverge.com19 hours agoView details
Microsoft AI CEO says AI threats are real, and Anthropic is making it worse
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that Mustafa has strong opinions on how AI should be built and regulated. Microsoft just published a 37-…
theverge.com19 hours agoView details
Inside the suddenly explosive world of AI safety
On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue…
theverge.com21 hours agoView details
Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map co…
arxiv.org5 hours agoView details
arXiv:2609.19180v1 Announce Type: new Abstract: Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a bi…
arxiv.org5 hours agoView details
What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
arXiv:2609.19182v1 Announce Type: new Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what res…
arxiv.org5 hours agoView details
Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction
arXiv:2609.19606v1 Announce Type: new Abstract: This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model's chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not.…
arxiv.org5 hours agoView details
FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
arXiv:2609.19680v1 Announce Type: new Abstract: Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity,…
arxiv.org5 hours agoView details
LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
arXiv:2609.19721v1 Announce Type: new Abstract: Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that convert…
arxiv.org5 hours agoView details
Can Data Attribution Filter Out Subliminal Learning? Not Reliably
arXiv:2609.20027v1 Announce Type: new Abstract: Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it id…
arxiv.org5 hours agoView details
arXiv:2609.20565v1 Announce Type: new Abstract: Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client's real-time…
arxiv.org5 hours agoView details
Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition
arXiv:2609.19148v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) are affective states in which individuals express contradictory signals across facial, vocal, and linguistic channels. Automatically recognising A/H in clinical videos requires detecting cross-modal disagreement -- the signal that standard…
arxiv.org5 hours agoView details
Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds
arXiv:2609.19149v1 Announce Type: new Abstract: Subliminal learning shows that language models can transmit a hidden trait through outputs that appear unrelated to it. One proposed explanation, token entanglement, links animal and number tokens through the model's output vocabulary. Yet existing measurements answer di…
arxiv.org5 hours agoView details
arXiv:2609.19150v1 Announce Type: new Abstract: Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative:…
arxiv.org5 hours agoView details
arXiv:2609.19151v1 Announce Type: new Abstract: Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applic…
arxiv.org5 hours agoView details
FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool
arXiv:2609.19152v1 Announce Type: new Abstract: Misinformation detection tools often rely on binary true and false classifications or models trained on historical examples, limiting their usefulness when novel misleading narratives emerge. Here, we present FakeSpotter, a content- and strategy-agnostic tool designed to…
arxiv.org5 hours agoView details
Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
arXiv:2609.19203v1 Announce Type: new Abstract: AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, bu…
arxiv.org5 hours agoView details
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
arXiv:2609.19212v1 Announce Type: new Abstract: Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as approximately lin…
arxiv.org5 hours agoView details
arXiv:2609.19244v1 Announce Type: new Abstract: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-…
arxiv.org5 hours agoView details
Do AI Agents Understand Computer Architecture?
arXiv:2609.19387v1 Announce Type: new Abstract: Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently ove…
arxiv.org5 hours agoView details
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failur…
arxiv.org5 hours agoView details
Closed-World Resolution Against Tool Hallucination in LLM Agents
arXiv:2609.19425v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agen…
arxiv.org5 hours agoView details
Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data
arXiv:2609.19153v1 Announce Type: new Abstract: Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines -- TF-IDF features and linear classifiers -- because the textual feature is often the object of study, not merely a means to a prediction.…
arxiv.org5 hours agoView details
Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry
arXiv:2609.19154v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we intro…
arxiv.org5 hours agoView details
Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue
arXiv:2609.19155v1 Announce Type: new Abstract: In Human-LLM dialogue, follow-up user utterances may implicitly conflict with earlier intents, leading the LLM to misinterpret user needs and generate inappropriate responses. A reliable dialogue system should proactively detect user-side conflicts before generating a re…
arxiv.org5 hours agoView details
Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes
arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajectories suffer from a Scaling Collapse: whe…
arxiv.org5 hours agoView details
VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering
arXiv:2609.19158v1 Announce Type: new Abstract: Knowledge graphs are usually integrated into question answering by encoding a retrieved subgraph with a graph neural network and fusing it with the language model in the online inference path. The same subgraph is therefore re-encoded from scratch every time a pair is sc…
arxiv.org5 hours agoView details
arXiv:2609.19164v1 Announce Type: new Abstract: In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should b…
arxiv.org5 hours agoView details
arXiv:2609.19167v1 Announce Type: new Abstract: As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences. Progress here is bottlenecked by evaluat…
arxiv.org5 hours agoView details
arXiv:2609.19238v1 Announce Type: new Abstract: This paper describes the participation of the YNU-HPCC team in subtask A of task 11, Bridging the Gap in Text-Based Emotion at SemEval-2025. Our best-performing system employs the RoBERTa (Robustly Optimized BERT Approach) model, an improved version of BERT that utilizes…
arxiv.org5 hours agoView details
Why Pretraining Fails to Share Cross-Lingual Knowledge
arXiv:2609.19291v1 Announce Type: new Abstract: Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during…
arxiv.org5 hours agoView details
The syntax and semantics of goals
arXiv:2609.19448v1 Announce Type: new Abstract: In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals are compositional representations whose content relates to rational beh…
arxiv.org5 hours agoView details
Compositional Reasoning in Language Models under Reinforcement Learning Post-Training
arXiv:2609.19465v1 Announce Type: new Abstract: Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the…
arxiv.org5 hours agoView details
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models…
arxiv.org5 hours agoView details
arXiv:2609.19513v1 Announce Type: new Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the op…
arxiv.org5 hours agoView details
LLM-as-an-Improver: Turning Verification into Better Candidates
arXiv:2609.19515v1 Announce Type: new Abstract: Verifier-based selection improves LLM performance by generating multiple candidate solutions and using a verifier to select the most promising one. However, existing methods typically treat verification only as a ranking step and discard its feedback once a fixed candida…
arxiv.org5 hours agoView details
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
arXiv:2609.19519v1 Announce Type: new Abstract: Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue t…
arxiv.org5 hours agoView details
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
arXiv:2609.19523v1 Announce Type: new Abstract: Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard…
arxiv.org5 hours agoView details
arXiv:2609.19524v1 Announce Type: new Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation eviden…
arxiv.org5 hours agoView details
Self Improvement via Fast Tree-search
arXiv:2609.19526v1 Announce Type: new Abstract: Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-…
arxiv.org5 hours agoView details
A frontend-backend architecture for tool calls in full-duplex speech models
arXiv:2609.19334v1 Announce Type: new Abstract: Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend le…
arxiv.org5 hours agoView details
The Role of Fine-grained Harm Signals in LLM Safety
arXiv:2609.19366v1 Announce Type: new Abstract: Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm…
arxiv.org5 hours agoView details
Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion
arXiv:2609.19417v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a graph-free multimodal framework that int…
arxiv.org5 hours agoView details
For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances
arXiv:2609.19504v1 Announce Type: new Abstract: As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared…
arxiv.org5 hours agoView details
From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models
arXiv:2609.19553v1 Announce Type: new Abstract: Model fusion integrates the capabilities from source models into a single target model. As of June 2026, Hugging Face hosts more than 2M models. This growing pool provides a rich base for model reuse and capability integration. Yet existing surveys often cover only separ…
arxiv.org5 hours agoView details
arXiv:2609.19585v1 Announce Type: new Abstract: In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making tempor…
arxiv.org5 hours agoView details
Form Over Content In Gradient-Based Data Attribution Methods
arXiv:2609.19589v1 Announce Type: new Abstract: Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports th…
arxiv.org5 hours agoView details
Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
arXiv:2609.19596v1 Announce Type: new Abstract: Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or…
arxiv.org5 hours agoView details
arXiv:2609.19530v1 Announce Type: new Abstract: Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet r\'esum\'e screening, the first gate, is commonly automated as a static, one-call judgment over a r\'esum\'e-job pair. We study a two-agent alternative in…
arxiv.org5 hours agoView details
Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks
arXiv:2609.19538v1 Announce Type: new Abstract: Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling among mobility, connectivity, and…
arxiv.org5 hours agoView details
Continual Enterprise World Model Discovery in Dynamic Systems
arXiv:2609.19551v1 Announce Type: new Abstract: In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system…
arxiv.org5 hours agoView details
SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership
arXiv:2609.19610v1 Announce Type: new Abstract: Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual…
arxiv.org5 hours agoView details
From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization
arXiv:2609.19630v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refu…
arxiv.org5 hours agoView details
Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs
arXiv:2609.19636v1 Announce Type: new Abstract: Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier ac…
arxiv.org5 hours agoView details
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
arXiv:2609.19644v1 Announce Type: new Abstract: Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI ind…
arxiv.org5 hours agoView details
arXiv:2609.19654v1 Announce Type: new Abstract: Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent r…
arxiv.org5 hours agoView details
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
arXiv:2609.19671v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, tra…
arxiv.org5 hours agoView details
Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction
arXiv:2609.19615v1 Announce Type: new Abstract: Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic disc…
arxiv.org5 hours agoView details
arXiv:2609.19650v1 Announce Type: new Abstract: Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer i…
arxiv.org5 hours agoView details
arXiv:2609.19736v1 Announce Type: new Abstract: This paper proposes a phonemically comprehensive, ASCII-only romanization scheme for Thai and Lao, treating the two closely related languages as a unified cross-lingual design problem. The scheme represents segmental contrasts, vowel length, and lexical tone while mainta…
arxiv.org5 hours agoView details
arXiv:2609.19778v1 Announce Type: new Abstract: Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly…
arxiv.org5 hours agoView details
Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
arXiv:2609.19799v1 Announce Type: new Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evol…
arxiv.org5 hours agoView details
Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data
arXiv:2609.19805v1 Announce Type: new Abstract: Grapheme-to-phoneme (G2P) conversion turns raw text into its phonemic form and is an essential part of both text-to-speech (TTS) and automatic speech recognition (ASR) systems. It is required to be fast, stable and context-aware. For unsegmented languages such as Japanes…
arxiv.org5 hours agoView details
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
arXiv:2609.19827v1 Announce Type: new Abstract: With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, informati…
arxiv.org5 hours agoView details
arXiv:2609.19868v1 Announce Type: new Abstract: Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent…
arxiv.org5 hours agoView details
JustMem: Just-Enough Memory Access for Long-Term Conversations
arXiv:2609.19877v1 Announce Type: new Abstract: Efficient long-term conversational memory requires retrieving sufficient evidence without indiscriminately expanding the context presented to the language model. This is challenging because relevant evidence may be distributed across multiple sessions, while compression…
arxiv.org5 hours agoView details
V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
arXiv:2609.19879v1 Announce Type: new Abstract: Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in th…
arxiv.org5 hours agoView details
D-Quant: Driftable Entropy Coding for KV Cache Quantization
arXiv:2609.19880v1 Announce Type: new Abstract: The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is p…
arxiv.org5 hours agoView details
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
arXiv:2609.19883v1 Announce Type: new Abstract: Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM…
arxiv.org5 hours agoView details
Evaluating Communicative Success in Machine-Translated Conversation
arXiv:2609.19885v1 Announce Type: new Abstract: Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds…
arxiv.org5 hours agoView details
Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words
arXiv:2609.19887v1 Announce Type: new Abstract: Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient information for the given context. Do pretr…
arxiv.org5 hours agoView details
AutoData: Agentic Search for Pre-training Data Selection
arXiv:2609.19754v1 Announce Type: new Abstract: LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic enginee…
arxiv.org5 hours agoView details
Rethinking Multi-Agent Collaboration: When More Is Less
arXiv:2609.19759v1 Announce Type: new Abstract: The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent…
arxiv.org5 hours agoView details
TorchCraft: Unified binder design by inverting an all-atom structure predictor
arXiv:2609.19770v1 Announce Type: new Abstract: All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom pre…
arxiv.org5 hours agoView details
arXiv:2609.19775v1 Announce Type: new Abstract: Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical knowledge in millions of case reports in the public medicine literature…
arxiv.org5 hours agoView details
KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
arXiv:2609.19916v1 Announce Type: new Abstract: Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage…
arxiv.org5 hours agoView details
arXiv:2609.19942v1 Announce Type: new Abstract: In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have litt…
arxiv.org5 hours agoView details
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence
arXiv:2609.19965v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect…
arxiv.org5 hours agoView details
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
arXiv:2609.19969v1 Announce Type: new Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and…
arxiv.org5 hours agoView details
Benchmarking LLM Compliance with China AI Generated Content Regulations
arXiv:2609.19989v1 Announce Type: new Abstract: The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content…
arxiv.org5 hours agoView details
Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
arXiv:2609.20081v1 Announce Type: new Abstract: SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that…
arxiv.org5 hours agoView details
Design of the IBM Granite 5.0 TurboCTC ASR Model
arXiv:2609.20104v1 Announce Type: new Abstract: We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using st…
arxiv.org5 hours agoView details
Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems
arXiv:2609.19789v1 Announce Type: new Abstract: Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter…
arxiv.org5 hours agoView details
Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy
arXiv:2609.19820v1 Announce Type: new Abstract: Regularized self-play -- the family behind DeepNash's Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy $\rho$. When the game has a polytope of value-equivalent equilibri…
arxiv.org5 hours agoView details
arXiv:2609.19830v1 Announce Type: new Abstract: Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Att…
arxiv.org5 hours agoView details
MetaRTL: Meta-path Attention Enhanced Relational Table Learning
arXiv:2609.19832v1 Announce Type: new Abstract: Relational table learning has gained increasing attention with the widespread use of relational databases. Existing methods typically rely on deep GNN or HGNN stacks, leading to high computational costs and limited performance on large real-world databases. We propose Me…
arxiv.org5 hours agoView details
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
arXiv:2609.19843v1 Announce Type: new Abstract: LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural bi…
arxiv.org5 hours agoView details
arXiv:2609.19848v1 Announce Type: new Abstract: Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local placement utility, and stability across camera motion. Accessibility and multilingual requirements further change label dimensions, yet algorithmic evaluations often…
arxiv.org5 hours agoView details
Reproducibility is not construct validity: LLM measurement of institutionally situated communication
arXiv:2609.19866v1 Announce Type: new Abstract: High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to f…
arxiv.org5 hours agoView details
arXiv:2609.19871v1 Announce Type: new Abstract: Time series forecasting has seen signicant advancements with the emergence of new deep learning models. However, forecasting time series in applications involving physical processes remains a major challenge. Despite the apparition of Physics Informed Neural Networks (PI…
arxiv.org5 hours agoView details
TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives
arXiv:2609.19897v1 Announce Type: new Abstract: Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional use. We introduce TRACE, a traini…
arxiv.org5 hours agoView details
arXiv:2609.19928v1 Announce Type: new Abstract: Per-user LLM inference on transaction histories binds the inference budget linearly to user count, which becomes prohibitive at applied scale. We re-cast attribute inference from per-user to per-transaction-pattern. The pipeline runs in three phases: Resolve abstracts it…
arxiv.org5 hours agoView details
Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models
arXiv:2609.19934v1 Announce Type: new Abstract: Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: t…
arxiv.org5 hours agoView details
MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation
arXiv:2609.19944v1 Announce Type: new Abstract: Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objective. We propose MaSCoD, a multi-agent framework that organizes candida…
arxiv.org5 hours agoView details
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
arXiv:2609.19947v1 Announce Type: new Abstract: LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resour…
arxiv.org5 hours agoView details
Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs
arXiv:2609.19961v1 Announce Type: new Abstract: Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucinatio…
arxiv.org5 hours agoView details
arXiv:2609.19996v1 Announce Type: new Abstract: With the widespread use of online navigation and ride-hailing services, achieving optimal route planning for diverse user preferences has recently attracted increasing attention. Classic graph algorithms for pathfinding use heuristic cost functions to define edge weight,…
arxiv.org5 hours agoView details
E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews
arXiv:2609.20001v1 Announce Type: new Abstract: Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and…
arxiv.org5 hours agoView details
Geopolitical Divisions Across Languages in Large Language Models
arXiv:2609.20005v1 Announce Type: new Abstract: People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukrai…
arxiv.org5 hours agoView details
FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction
arXiv:2609.20026v1 Announce Type: new Abstract: Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatial-temporal approaches. In such federated settings, each client observes traffic…
arxiv.org5 hours agoView details
arXiv:2609.20051v1 Announce Type: new Abstract: Step distillation reduces the cost of video generation, but reusing a LoRA trained for a longer trajectory can alter its functional effect or degrade target quality. Static parameter compatibility offers one perspective on this problem; our observations show that similar…
arxiv.org5 hours agoView details
MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution
arXiv:2609.20056v1 Announce Type: new Abstract: Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision…
arxiv.org5 hours agoView details
arXiv:2609.20057v1 Announce Type: new Abstract: Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic paths, and type constraint violations. The manual curation of these issues…
arxiv.org5 hours agoView details
arXiv:2609.20067v1 Announce Type: new Abstract: Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based methods (LIME, SHAP) are categorically inapplicable to this purpose, as the…
arxiv.org5 hours agoView details
arXiv:2609.20068v1 Announce Type: new Abstract: This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a dimin…
arxiv.org5 hours agoView details
Tailored to you: longitudinal effects of personalising language models
arXiv:2609.20077v1 Announce Type: new Abstract: Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward…
arxiv.org5 hours agoView details
arXiv:2609.20080v1 Announce Type: new Abstract: The growing complexity of multi-domain operational environments (land, aerospace, naval, cyber, and electromagnetic spectrum) has increased the volume and velocity of data reaching command-and-control (C2) centers, straining the observe-orient-decide-act (OODA) decision…
arxiv.org5 hours agoView details
Fine-Tuning Models for Biomedical Relation Extraction
arXiv:2609.20169v1 Announce Type: new Abstract: Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical literature remains a complex challenge th…
arxiv.org5 hours agoView details
To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals
arXiv:2609.20186v1 Announce Type: new Abstract: Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust perfo…
arxiv.org5 hours agoView details
arXiv:2609.20207v1 Announce Type: new Abstract: Large language models produce prompt-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives. We develop an observable framework for determining when language-derived probabilities…
arxiv.org5 hours agoView details
Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech
arXiv:2609.20223v1 Announce Type: new Abstract: We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlap…
arxiv.org5 hours agoView details
arXiv:2609.20232v1 Announce Type: new Abstract: We introduce the \textbf{Public Discourse Corpus (PDC)}, the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across seven professional domains, yielding 186,6…
arxiv.org5 hours agoView details
UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning
arXiv:2609.20089v1 Announce Type: new Abstract: Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerg…
arxiv.org5 hours agoView details
arXiv:2609.20252v1 Announce Type: new Abstract: High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity than the massive corpora used to pretra…
arxiv.org5 hours agoView details
arXiv:2609.20303v1 Announce Type: new Abstract: Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must…
arxiv.org5 hours agoView details
Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering
arXiv:2609.20398v1 Announce Type: new Abstract: Semantic parsing (SP)-based knowledge base question answering aims to answer natural language questions by generating executable logical forms (LFs) over knowledge bases (KBs). When applying Large Language Models (LLMs) to this task, a key challenge over large, heterogen…
arxiv.org5 hours agoView details
Xeno-Interpretability: Investigating the Alien Minds of LLMs
arXiv:2609.20408v1 Announce Type: new Abstract: Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate hu…
arxiv.org5 hours agoView details
Stress-testing Alignment Midtraining
arXiv:2609.20412v1 Announce Type: new Abstract: When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One p…
arxiv.org5 hours agoView details
Edustories: A Collection of Real-world Case Studies from Classroom Practices
arXiv:2609.20484v1 Announce Type: new Abstract: Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI…
arxiv.org5 hours agoView details
Relational Attention for Data-Efficient Language Modeling
arXiv:2609.20530v1 Announce Type: new Abstract: We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), whic…
arxiv.org5 hours agoView details
An Analysis of Training-Free Self-Reported Confidence in Language Models
arXiv:2609.20541v1 Announce Type: new Abstract: Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, an…
arxiv.org5 hours agoView details
SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment
arXiv:2609.20584v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial b…
arxiv.org5 hours agoView details
WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution
arXiv:2609.20593v1 Announce Type: new Abstract: Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory th…
arxiv.org5 hours agoView details
Models
View all models →KaLM-Embedding/KaLM-Reranker-V1-Nano
text-ranking · sentence-transformers · safetensors · t5gemma2
huggingface.co1 day ago10 ptsView details
AMAImedia/DeepSeek-V4.1-Flash-FP8-GGUF
image-text-to-text · transformers · safetensors · gguf
huggingface.co1 day ago1 ptsView details
AMAImedia/DeepSeek-V4-Flash-Vision-Exp-FP8-GGUF
image-text-to-text · transformers · safetensors · gguf
huggingface.co1 day agoView details
AMAImedia/GLM-5.3-Flash-FP8-GGUF
image-text-to-text · transformers · safetensors · gguf
huggingface.co1 day agoView details
AMAImedia/Qwen3.8-Flash-Next-BF16-GGUF
image-text-to-text · transformers · safetensors · gguf
huggingface.co1 day agoView details
AMAImedia/DeepSeek-V4-Flash-0731
image-text-to-text · transformers · safetensors · deepseek_v4
huggingface.co1 day agoView details
mlboydaisuke/MiniCPM5-2B-CoreAI
text-generation · coreai · coreai-aimodel · core-ai
huggingface.co1 day agoView details
Open source
View all open source →## [3.15.0](https://github.com/openai/openai-python/compare/v3.14.1...v3.15.0) (2026-09-18) ### Features * **api:** add agent session model settings ([#3882](https://github.com/openai/openai-python/issues/3882)) ([4b15817](https://github.com/openai/openai-python/commit/4b1581771…
github.com8 hours agoView details
langchain-ai/langchain langchain-typesafe==0.0.1a2
Initial release release(typesafe): bump to 0.0.1a2 (#40576) feat(typesafe): experimental `AutoModeMiddleware` (#40545) feat(typesafe): experimental `ModelRouterMiddleware` (#40543) fix(typesafe): trace usage metadata (#40570) feat(typesafe): `TypeSafeClassifier` (#40542)
github.com10 hours agoView details
langchain-ai/langchain langchain-typesafe==0.0.1a1
Initial release feat(typesafe): `TypeSafeClassifier` (#40542)
github.com15 hours agoView details
<details open> ci : add missing evict-old-files (#29041) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/48252257> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/…
github.com15 hours agoView details
<details open> opencl: fix various warnings (#28984) * opencl: fix warnings * opencl: fix warnings for non adreno </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/48132480> **macOS/iOS:** - [macOS Apple Silicon…
github.com1 day agoView details
<details open> vulkan: fix buffer_reference alignment in im2col shaders (#28996) Both im2col.comp and im2col_3d.comp declare D_ptr without an explicit buffer_reference_align, so glslang emits writes through it as Aligned 16. The shaders advance the pointer by D_SIZE, a per-varia…
github.com1 day agoView details
<details open> vulkan: support qwen4exp hc ops (#28988) * vulkan: support qwen4exp hc ops * fix stale comment [no-ci] </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/48109146> **macOS/iOS:** - [macOS Apple Sil…
github.com1 day agoView details