Archive / 2026-09-15
September 15, 2026
News
View all news →Introducing System One Models and Jev
typesafe.ai2 days ago1818 ptsView detailsJoin discussion
Why I'm still bearish on LLMs after Navier-Stokes
dank.systems2 days ago486 ptsView detailsJoin discussion
Gemini 3.8 Live and 3.8 Live Extended Thinking
blog.google2 days ago482 ptsView detailsJoin discussion
Suspected sabotage causes major Netherlands rail disruption
bbc.com2 days ago481 ptsView detailsJoin discussion
We got admin access to Baseten's production GitHub
strix.ai2 days ago321 ptsView detailsJoin discussion
Let's make quality the norm again
forbrukerradet.no2 days ago377 ptsView detailsJoin discussion
Show HN: Capsule – Single-file web apps that save their data into SQLite
withcapsule.app2 days ago323 ptsView detailsJoin discussion
There's a 100% Chance AI Agents Are Ruining the Internet
404media.co2 days ago230 ptsView detailsJoin discussion
vale.rocks3 days ago267 ptsView detailsJoin discussion
Global bond yields hit 2008 highs, raising stakes for big borrowers
reuters.com2 days ago161 ptsView detailsJoin discussion
Learning to solve hard problems in RL for LLMs by never giving up
mnoukhov.github.io2 days ago117 ptsView detailsJoin discussion
How much of F-Droid is LLM generated?
tintotint.eu2 days ago145 ptsView detailsJoin discussion
A Cop Searched 19,000 Flock Cameras Across 1,558 Cities. His Reason: 'LMAO'
techtimes.co.uk2 days ago122 ptsView detailsJoin discussion
Stay discoverable in search while disallowing AI training
blog.cloudflare.com2 days ago84 ptsView detailsJoin discussion
Australia 'on the same page' as Canada as it seeks deeper EU alliance
reddit.com2 days ago120 ptsView detailsJoin discussion
Cartesian – AI 3D Modeling for Design
formas.ai2 days ago107 ptsView detailsJoin discussion
Datamimic – don't let your coding agent invent its own test world
github.com2 days ago58 ptsView detailsJoin discussion
AI is breaking our proxies for expertise
seangoedecke.com2 days ago86 ptsView detailsJoin discussion
Our chance to take down the Epstein class
thomasmassie.org2 days ago62 ptsView detailsJoin discussion
Show HN: Pizza Bot – An inbox for AI agents that work in the background
github.com2 days ago59 ptsView detailsJoin discussion
Show HN: Ordewell – turn one goal into an ordered plan of coding-agent tasks
github.com2 days ago54 ptsView detailsJoin discussion
The Age of Wonders and Terrors
scottaaronson.blog2 days ago47 ptsView detailsJoin discussion
Show HN: Panel – A research workspace where the agent can build its own panes
github.com2 days ago53 ptsView detailsJoin discussion
What we have learned at OpenShell applying formal methods to control AI agents
nvidia.github.io2 days ago39 ptsView detailsJoin discussion
Show HN: The bottom 50% of U.S. households are short after essentials (BLS data)
whats-left-over.pages.dev2 days ago38 ptsView detailsJoin discussion
Show HN: Loss. a tiny satire about AI progress
workatloss.com2 days ago38 ptsView detailsJoin discussion
1Password's AI patching benchmark is misleading
blog.trailofbits.com2 days ago37 ptsView detailsJoin discussion
Charges Against Man Who Destroyed 3D-Printed 'Decoy' Flock Camera Reduced
404media.co2 days ago27 ptsView detailsJoin discussion
AI Regulation as Anthropic's Business Model
twitter.com2 days ago27 ptsView detailsJoin discussion
The Apple IIGS was introduced 40 years ago today
twitter.com2 days ago16 ptsView detailsJoin discussion
Garbage Trucks Now Have AI Cameras to Score Your House and Clock Code Violations
thedrive.com2 days ago21 ptsView detailsJoin discussion
GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt
arxiv.org2 days ago23 ptsView detailsJoin discussion
AI models in 'surreal' dialect mixing poetic language and tech bro jargon
theguardian.com2 days ago20 ptsView detailsJoin discussion
Mark Zuckerberg: Labs should self-regulate
twitter.com2 days ago15 ptsView detailsJoin discussion
Deep Seek v4.1 M5 Max at 17 tokens/s
github.com2 days ago13 ptsView detailsJoin discussion
Jev: The Model That Gives AI the Properties of Code
twitter.com2 days ago17 ptsView detailsJoin discussion
sufficientlyadvanced.blog2 days ago12 ptsView detailsJoin discussion
How AI tool calling works (40 lines of vanilla JavaScript)
buttercup.sh2 days ago15 ptsView detailsJoin discussion
Experience Necessary. The Irony of Building a Skilled Team in the Age of AI
benversh.substack.com2 days ago15 ptsView detailsJoin discussion
Auto-autoresearch: self-improving agents on Karpathy's NanoChat benchmark
rekursiv.ai2 days ago13 ptsView detailsJoin discussion
California attempts to crack down on health spending
axios.com2 days ago12 ptsView detailsJoin discussion
Show HN: Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s
github.com2 days ago13 ptsView detailsJoin discussion
With Apple Watch that's always listening, what's at stake is public perception
manualdousuario.net2 days ago14 ptsView detailsJoin discussion
Anthropic and OpenAI look to Uncle Sam to make them too big to fail
theregister.com2 days ago12 ptsView detailsJoin discussion
V1.1 state of open source- OS 4.4 months behind frontier [pdf]
stateofopensource.ai2 days ago12 ptsView detailsJoin discussion
Cockroach Continuum – Elastic Infrastructure for Agentic Database Estates
cockroachlabs.com2 days ago12 ptsView detailsJoin discussion
Open weights are not open source: Why AI's favorite label is under dispute
theregister.com2 days ago13 ptsView detailsJoin discussion
The bitter lesson of browser agents
browser-use.com2 days ago12 ptsView detailsJoin discussion
Show HN: farseer.space – fly anywhere in the universe in your browser
farseer.space2 days ago10 ptsView detailsJoin discussion
NATO jets shoot down drone violating Lithuania's airspace
theguardian.com3 days ago13 ptsView detailsJoin discussion
polylane.com2 days ago10 ptsView detailsJoin discussion
Show HN: TabPFN-3.5, a Tabular Foundation Model for messy real-world tables
priorlabs.ai2 days ago10 ptsView detailsJoin discussion
Show HN: Bough, the agent I built to replace Claude Code at work
github.com2 days ago10 ptsView detailsJoin discussion
- Primary source
Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
Algorithms & Theory
research.google2 days agoView details
- Primary source
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
deepmind.google2 days agoView details
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-learn interface. It posts the top TabArena Elo among single models, ahead of Google's TabFM and LG's EXAONE Tabular. The code is Apache-2.0, while weights are licensed for non-commercial research. The po…
marktechpost.com2 days agoView details
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations, FP8-style epilogues, scaled dot-product attention, dynamic shapes, and CUDA graph captures. This practical tutorial demonstrates how to optimize deep learning computations directly below framework…
marktechpost.com2 days agoView details
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. The models execute tools and API calls in the background while the conversation keeps flowing, process live visual inputs, and switch between 97 languages mid conversation. Extended Thinking ranks…
marktechpost.com2 days agoView details
AI and data centers are incredibly unpopular in every poll
Poll data released Tuesday by The New York Times and Siena University confirms what we've already been seeing, and what politicians are responding to - AI and data centers are incredibly unpopular. Asked if they support or oppose the construction of data centers to power AI tech, 61 percent of the 1,503 likely voters…
theverge.com2 days agoView details
AI ‘Actor’ Tilly Norwood Told Me That ‘All Lives Matter’
The virtual character, which is promoting its upcoming movie Misaligned, tries to evade politics by repetitively commenting on the clothes you’re wearing.
wired.com2 days agoView details
Meta’s new One subscriptions put a price on social media and AI
Shortly after launching its new do-everything AI assistant Muse, Meta's launching subscription bundles that pair its standalone app subscriptions with extra AI usage. Some of the new Meta One bundles were in testing earlier this year, but are now available globally starting today, with several tiers for individual use…
theverge.com2 days agoView details
This doorbell camera lets a human security guard watch your front door
The new SimpliSafe Video Doorbell Series 2 adds 2K resolution and dual band Wi-Fi. | Image: Simplisafe DIY home security company SimpliSafe is bringing its AI-powered proactive security feature to the front door. The new SimpliSafe Video Doorbell Series 2 launches today for $199.99 and works with the company's Active…
theverge.com2 days agoView details
How Humans and LLMs Read Gender into Gender-Neutral Physical Descriptions
arXiv:2609.16366v1 Announce Type: new Abstract: When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "she", "his") in favor of seemingly "objective" physical descriptions (e.g., "short hair", "a defined jawline"). Yet whether…
arxiv.org2 days agoView details
Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data
arXiv:2609.16532v1 Announce Type: new Abstract: Continued pretraining (CPT) with data augmentation such as paraphrasing can store inside a large language model (LLM) the knowledge of a small source corpus. The stored knowledge, however, is not always retrieved correctly. We study the eliciting side rather than the sto…
arxiv.org2 days agoView details
RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue
arXiv:2609.16614v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, es…
arxiv.org2 days agoView details
arXiv:2609.16900v1 Announce Type: new Abstract: Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchma…
arxiv.org2 days agoView details
Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech
arXiv:2609.16906v1 Announce Type: new Abstract: Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies tha…
arxiv.org2 days agoView details
Lit3R: Retrieve-Relate-Read for Evidence-Grounded Question Answering over Scientific Literature
arXiv:2609.16912v1 Announce Type: new Abstract: We describe tus-nlp's Lit3R (Retrieve-Relate-Read) system for LitTraceQA, a shared task for literature-grounded question answering that requires systems to retrieve relevant papers, identify supporting evidence, and generate answers. Lit3R combines off-the-shelf retrieva…
arxiv.org2 days agoView details
arXiv:2609.16993v1 Announce Type: new Abstract: Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. Ho…
arxiv.org2 days agoView details
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
arXiv:2609.16995v1 Announce Type: new Abstract: Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and…
arxiv.org2 days agoView details
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
arXiv:2609.17043v1 Announce Type: new Abstract: Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents. Whether this holds at…
arxiv.org2 days agoView details
Latent Undertow: How Ordinary Typos Break Probes
arXiv:2609.15994v1 Announce Type: new Abstract: LLMs handle ordinary typing variation fluently: a typo or missing punctuation leaves both user intent and the model's response substantively unchanged. Yet probes that detect malicious prompts by reading the model's hidden states tell a different story: the same edit rot…
arxiv.org2 days agoView details
NepKANUN: A RAG-Based Nepali Legal Assistant
arXiv:2609.15999v1 Announce Type: new Abstract: Accessing legal information in Nepal is difficult due to complex terminology, limited resources, and misinformation. We introduce an AI-powered legal assistant that is tailored for Nepali legal texts and is built on a fine-tuned large language model. The technology provi…
arxiv.org2 days agoView details
Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT)
arXiv:2609.16010v1 Announce Type: new Abstract: The complexity of legal language and limited accessibility to legal information pose significant challenges to justice delivery in Nepal. Traditional legal services remain inaccessible to many citizens due to language barriers, information fragmentation, and a critical s…
arxiv.org2 days agoView details
PunGraph: Retrieval-Enhanced Phonetic-Semantic Graph Reasoning for Pun Understanding
arXiv:2609.16557v1 Announce Type: new Abstract: Puns are a challenging form of figurative language that exploit phonetic similarity and semantic ambiguity to convey multiple meanings. Although large language models (LLMs) demonstrate strong language understanding capabilities, they still struggle with pun reasoning du…
arxiv.org2 days agoView details
Register Tokens for Bounded-State Reasoning in Diffusion Language Models
arXiv:2609.16372v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continu…
arxiv.org2 days agoView details
ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
arXiv:2609.16393v1 Announce Type: new Abstract: We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label f…
arxiv.org2 days agoView details
Quantifying Organizational Environmental Action from Web Data and Large Language Models
arXiv:2609.16627v1 Announce Type: new Abstract: Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We…
arxiv.org2 days agoView details
Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
arXiv:2609.16660v1 Announce Type: new Abstract: Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidl…
arxiv.org2 days agoView details
Single Document Extractive Summarization using Domination in Hypergraph
arXiv:2609.15993v1 Announce Type: new Abstract: Automatic Text Summarization (ATS) in Natural Language Processing has been an important task in Information Retrieval. It compresses a document to create a summary that captures all the relevant and important information conveyed in the document. This study explores Hype…
arxiv.org2 days agoView details
arXiv:2609.16661v1 Announce Type: new Abstract: Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-…
arxiv.org2 days agoView details
arXiv:2609.17081v1 Announce Type: new Abstract: Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that h…
arxiv.org2 days agoView details
An Empirical Study of Counterfactual Self-Explanations in LLMs
arXiv:2609.17119v1 Announce Type: new Abstract: Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its…
arxiv.org2 days agoView details
A Data-free Universal Prior over Syntactic Structures
arXiv:2609.16854v1 Announce Type: new Abstract: Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories estimate the probability of syntactic structures from language-specific data. Whether part of this probability…
arxiv.org2 days agoView details
Target-Language Generation in Multilingual Models: Activation Steering and Optimal Control
arXiv:2609.16967v1 Announce Type: new Abstract: Ensuring that multilingual language models generate coherent text in a specific target language is a major issue in multilingual language modeling. We develop an optimal control method for target-language text generation as well as a framework for evaluating the quality…
arxiv.org2 days agoView details
arXiv:2609.15990v1 Announce Type: new Abstract: Few-shot prompting sometimes degrades language models instead of helping them, but why this happens is unknown. We evaluate 12 open-weight models on two Ukrainian tasks news classification and legal case outcome prediction and find that the effect is strongly task-depend…
arxiv.org2 days agoView details
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
arXiv:2609.15991v1 Announce Type: new Abstract: Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Function…
arxiv.org2 days agoView details
Optimal Model Activation Policies for Inference Networks of Large Language Models
arXiv:2609.15992v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have rendered them necessary for Natural Language Processing (NLP) tasks, and their high inference cost motivates the study of cost-performance trade-offs. In practice, several expert LLMs are used in synergy for inference,…
arxiv.org2 days agoView details
arXiv:2609.15995v1 Announce Type: new Abstract: Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit…
arxiv.org2 days agoView details
arXiv:2609.15996v1 Announce Type: new Abstract: Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM baseline consists of one-shot proposals. If bri…
arxiv.org2 days agoView details
arXiv:2609.15997v1 Announce Type: new Abstract: Improving safety at intersections requires identifying crash mechanisms and recommending appropriate countermeasures. However, this process traditionally relies on expert judgment, making it labor-intensive, difficult to scale, and dependent on the availability of experi…
arxiv.org2 days agoView details
Self-reported archetypes and behavioral failures in Large Language Models
arXiv:2609.15998v1 Announce Type: new Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet th…
arxiv.org2 days agoView details
ViCo: Visual-oriented Coding with Self-Reflection for Chart Replication
arXiv:2609.16014v1 Announce Type: new Abstract: This paper addresses the challenge of generating high-quality academic charts that match the visual standards of human-authored papers. While existing AI agents can produce well-structured text and code, their generated visualizations often lack the stylistic and semanti…
arxiv.org2 days agoView details
Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks
arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the dominant scalable alternative. We ask what…
arxiv.org2 days agoView details
Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents
arXiv:2609.16053v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents operating over extended interactions. Existing memory systems primarily update memory when new information arrives, treating retrieval as the endpoint of memory access rather than a driver of memory evolution. Consequent…
arxiv.org2 days agoView details
State of Thought Enables Endogenous Reasoning
arXiv:2609.16055v1 Announce Type: new Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansio…
arxiv.org2 days agoView details
Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation
arXiv:2609.16059v1 Announce Type: new Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often leads to surface-level pattern matching and degrades general capabilities. While Reinforcement…
arxiv.org2 days agoView details
Efficient Multimodal Generative Recommendation with Latent Narrative Reasoning
arXiv:2609.16070v1 Announce Type: new Abstract: Generative recommendation reformulates item prediction as semantic identifier generation, yet episodic content introduces a fundamentally different setting where the target is determined by narrative evolution rather than user preference. This task requires models to und…
arxiv.org2 days agoView details
The Immutable Past: Formalizing State Mutability and Conflict Resolution in Mutable RAG
arXiv:2609.16073v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) serves as the primary memory architecture for long-horizon autonomous agents. However, treating shared memory as an append-only stream introduces \textit{Semantic Shadowing}, a critical failure mode where conflicting historical observ…
arxiv.org2 days agoView details
arXiv:2609.16076v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework t…
arxiv.org2 days agoView details
arXiv:2609.16095v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for improving the quality of generated contents of Large Language Models (LLMs) by grounding responses in external knowledge, thus reducing hallucinations and factual errors. However, recent studies…
arxiv.org2 days agoView details
Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act
arXiv:2609.16268v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the tra…
arxiv.org2 days agoView details
Speaker-Specific and Language-Dependent Temporal Organization in Bilingual Political Speech
arXiv:2609.16274v1 Announce Type: new Abstract: Speech rhythm helps structure persuasive speech, but most empirical work examines monolingual English. This study asks how politicians organize timing when speaking Luxembourgish and French. We analyze 400 sentences from ten politicians, annotated for segments and pauses…
arxiv.org2 days agoView details
Speaker or Language? Explaining Variance in Charismatic Prosody Across Luxembourgish and French
arXiv:2609.16275v1 Announce Type: new Abstract: Charismatic speech is shaped by language and speaking style, yet their relative contribution in bilingual public speaking remains unclear. We analyzed spontaneous speeches of 10 politicians who address audiences in Luxembourgish and French, in highly comparable communica…
arxiv.org2 days agoView details
Efficient One-to-Many Translation with Joint Multi-Stream Diffusion
arXiv:2609.16312v1 Announce Type: new Abstract: One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number of target languages. We explore how diffusion can enable multilingual translation with a d…
arxiv.org2 days agoView details
StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
arXiv:2609.16340v1 Announce Type: new Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh post-edits for every new model is prohibi…
arxiv.org2 days agoView details
Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue
arXiv:2609.16396v1 Announce Type: new Abstract: Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior. We ask whether contexts centered on spoken negation cues contain measurable multimodal behavioral information: whether…
arxiv.org2 days agoView details
ReMova: Fine-tuning LLMs for English to Belarusian translation
arXiv:2609.16427v1 Announce Type: new Abstract: This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Bel…
arxiv.org2 days agoView details
Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling
arXiv:2609.16450v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative unmasking. However, dLLMs typically require many steps before token confidence reaches the decoding threshold, resulting…
arxiv.org2 days agoView details
arXiv:2609.16501v1 Announce Type: new Abstract: De-identified r\'esum\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-…
arxiv.org2 days agoView details
Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
arXiv:2609.16517v1 Announce Type: new Abstract: Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualificat…
arxiv.org2 days agoView details
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
arXiv:2609.16590v1 Announce Type: new Abstract: People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers…
arxiv.org2 days agoView details
arXiv:2609.16739v1 Announce Type: new Abstract: Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japan…
arxiv.org2 days agoView details
TIAO: Token Importance-Aware Policy Optimization for Text Summarization
arXiv:2609.16748v1 Announce Type: new Abstract: Text summarization requires models to condense content while preserving key qualities such as consistency and coherence. Large language models (LLMs) have shown strong performance on this task and can be further improved through reinforcement learning (RL). However, most…
arxiv.org2 days agoView details
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
arXiv:2609.16777v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critical safety concern. Existing red-teaming…
arxiv.org2 days agoView details
Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement
arXiv:2609.16800v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge. Existing memory-augmented approaches retrieve individual past examples as direct references, but do…
arxiv.org2 days agoView details
Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics
arXiv:2609.16860v1 Announce Type: new Abstract: Mandarin Chinese has two productive reduplicative constructions that repeat either two-character base words or their constituents (e.g., `in good health', `discuss a bit'). Their varied meanings have been described as realizing plurality, valence coloring, sound symbolis…
arxiv.org2 days agoView details
Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning
arXiv:2609.16890v1 Announce Type: new Abstract: Large Language Model (LLM) unlearning is essential for removing sensitive or copyrighted knowledge while preserving general utility. Existing methods often leave residual knowledge in intermediate representations, which can still be recovered. To address this, we propose…
arxiv.org2 days agoView details
arXiv:2609.16964v1 Announce Type: new Abstract: Rapid extraction of structured information from social media is important for humanitarian response, yet existing disaster tweet resources mainly provide document-level category labels without span-level entity annotations. We introduce HUMAID-NER, the first named entity…
arxiv.org2 days agoView details
arXiv:2609.16984v1 Announce Type: new Abstract: Open-weight language models publish the strings their chat templates use to mark turns, roles and tool results, which the tokenizer maps back to the reserved identifiers the model obeys. Anyone who controls text in a prompt can therefore write a turn boundary indistingui…
arxiv.org2 days agoView details
Autoformalizing Argumentative Material Inferences
arXiv:2609.16991v1 Announce Type: new Abstract: Natural language arguments are compelling before they are formally explicit. A premise supports a claim through defeasible warrants, background commitments, and exception conditions that the text leaves implicit. However, formal verification requires the opposite. Making…
arxiv.org2 days agoView details
Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising
arXiv:2609.16997v1 Announce Type: new Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis sentiment dataset with approximately 200K Facebook and YouTube comments…
arxiv.org2 days agoView details
Models
View all models →Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4
audio · music · yue2
huggingface.co2 days ago128 ptsView details
image-text-to-text · safetensors · zdtaichu5_0 · multimodal
huggingface.co3 days ago163 ptsView details
safetensors · glm_moe_dsa · arxiv:2609.15818
huggingface.co2 days ago129 ptsView details
mlx-community/gemma-4-12B-coder-fable5-composer2.5-v1-OptiQ-4bit
image-text-to-text · mlx · safetensors · gemma4_unified
huggingface.co3 days ago8 ptsView details
Justvugg/GLM-5.3-colibri-int4-g64
text-generation · colibri · glm_moe_dsa · moe
huggingface.co3 days ago3 ptsView details
Open source
View all open source →## [3.14.1](https://github.com/openai/openai-python/compare/v3.14.0...v3.14.1) (2026-09-15) ### Bug Fixes * **client:** validate retry limits and preserve application errors ([#3867](https://github.com/openai/openai-python/issues/3867)) ([f86c721](https://github.com/openai/opena…
github.com2 days agoView details
anthropics/anthropic-sdk-python v1.6.0
## 1.6.0 (2026-09-15) Full Changelog: [v1.5.0...v1.6.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.5.0...v1.6.0) ### Features * **api:** add auto mode tool permissions for Managed Agents ([909d92f](https://github.com/anthropics/anthropic-sdk-python/commit/909d…
github.com2 days agoView details
<details open> ci : fix android release (#28936) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47548995> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download…
github.com3 days agoView details