Archive / 2026-09-10
September 10, 2026
News
View all news →deepseek-ai/DeepSeek-V4.1-Flash
image-text-to-text · transformers · safetensors · deepseek_v41
huggingface.co8 days ago2911 ptsView details
twitter.com8 days ago968 ptsView detailsJoin discussion
I have a theory that software drives people insane
graybeard.ing7 days ago466 ptsView detailsJoin discussion
Mexican student creates an acoustic fire extinguisher to put out fire in seconds
upsocl.com7 days ago387 ptsView detailsJoin discussion
developers.openai.com7 days ago346 ptsView detailsJoin discussion
vester.si7 days ago335 ptsView detailsJoin discussion
Detecting and countering misuse of AI: September 2026
anthropic.com7 days ago173 ptsView detailsJoin discussion
Thelio Mira AI Linux Workstation: 192 GB GPU Memory
system76.com7 days ago119 ptsView detailsJoin discussion
AI Is Breaking This Thing We Call Trust
terriblesoftware.org7 days ago117 ptsView detailsJoin discussion
LRU is harder to beat than the KV-cache papers suggest
github.com7 days ago108 ptsView detailsJoin discussion
Bending Spoons buying Miro for $1.355B
investors.bendingspoons.com7 days ago103 ptsView detailsJoin discussion
1M+ German households have hung solar panels off their balcony railings
spacedaily.com7 days ago72 ptsView detailsJoin discussion
Mathematicians want proof OpenAI didn’t use their work
Sam Altman, chief executive officer of OpenAI, during a media tour of the Stargate AI data center. | Bloomberg via Getty Images Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company's…
theverge.com7 days ago77 ptsView details
Anthropic says it blocked possible efforts to build biological weapons
nytimes.com7 days ago68 ptsView detailsJoin discussion
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
tokenstead.ai7 days ago64 ptsView detailsJoin discussion
The Gemini app is now available for Windows
blog.google7 days ago55 ptsView detailsJoin discussion
OpenAI considers slowing advanced AI development, Sam Altman tells employees
bloomberg.com7 days ago56 ptsView detailsJoin discussion
The Four-Color Theorem Gets a Rare New Proof
quantamagazine.org7 days ago57 ptsView detailsJoin discussion
ai-2027.com7 days ago57 ptsView detailsJoin discussion
Show HN: MultiMatte, a Promptable Image Background Removal Model
usefeyn.com7 days ago54 ptsView detailsJoin discussion
How My Students Think About AI
lesswrong.com7 days ago46 ptsView detailsJoin discussion
One resignation turned the embers of AI fear into a wildfire
interconnects.ai7 days ago45 ptsView detailsJoin discussion
blog.kagi.com7 days ago43 ptsView detailsJoin discussion
Show HN: Syq – copy files between machines fast (better than rsync)
greaber.github.io7 days ago41 ptsView detailsJoin discussion
We Are Still Living in the Broken World Sept. 11 Created
bloomberg.com7 days ago33 ptsView detailsJoin discussion
Show HN: Biff 2.0 (Clojure web framework)
biffweb.com7 days ago28 ptsView detailsJoin discussion
AI Doomlord Jacob Coxon's Media Tour Has Begun
gizmodo.com7 days ago23 ptsView detailsJoin discussion
Stop externalizing the cost of your AI use to me
thelastsoftwareengineer.substack.com7 days ago17 ptsView detailsJoin discussion
Finding Slow Code with Wrapture
grahamdumpleton.me8 days ago19 ptsView detailsJoin discussion
AI researchers leave Anthropic and Google: 'There are no adults in the room'
nbcnews.com7 days ago16 ptsView detailsJoin discussion
heatherburns.tech7 days ago17 ptsView detailsJoin discussion
Streaming Is Raising Prices Faster Than Cable Ever Did
hollywoodreporter.com7 days ago17 ptsView detailsJoin discussion
Show HN: Open-source simulation testing infra for voice agents
github.com7 days ago16 ptsView detailsJoin discussion
Meta tried to shrink engineering teams around AI
leaddev.com7 days ago16 ptsView detailsJoin discussion
Brax 2026 Roadmap: secure mobile phone, smart wall display, pocket AI phone
youtube.com7 days ago16 ptsView detailsJoin discussion
andrewarrow.dev7 days ago15 ptsView detailsJoin discussion
OpenAI Wants to Know If an AI Industry Slowdown Would Even Be Legal
AI leaders worry antitrust law could stand in the way of what they view as an increasingly urgent push to coordinate a slowdown in AI development.
wired.com7 days ago13 ptsView detailsJoin discussion
LLM Visualizer – Build a Transformer from Scratch
jayvisaria.github.io7 days ago13 ptsView detailsJoin discussion
ChatGPT Pro 20x plan is now unavailable for purchase
twitter.com7 days ago13 ptsView detailsJoin discussion
What Happened at the White House That Marc Andreessen Says Sent Him to Trump
politico.com7 days ago12 ptsView detailsJoin discussion
Bending Spoons Agrees to Buy Miro for $1.36B
bloomberg.com7 days ago12 ptsView detailsJoin discussion
The case to BYOB: build your own (coding) benchmarks
byobench.ai7 days ago11 ptsView detailsJoin discussion
I Was Offered Money to Tell You AI Will Kill Us [video]
youtube.com8 days ago10 ptsView detailsJoin discussion
Show HN: Botbin.io – pastebin for AI agent artifacts
botbin.io8 days ago10 ptsView detailsJoin discussion
- Primary source
How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
César de la Fuente’s lab uses Codex and ChatGPT to search living and extinct genomes for antimicrobial candidates to fight drug-resistant infections.
openai.com7 days agoView details
- Primary source
Now everyone can put data to work
Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.
openai.com7 days agoView details
- Primary source
Introducing ChatGPT for Financial Services
Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.
openai.com8 days agoView details
- Primary source
Expanding AI access and cyber defense for federal, state, local, and tribal governments
OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.
openai.com8 days agoView details
- Primary source
ToolGrad: Efficient tool-use dataset generation with textual "gradients"
Machine Intelligence
research.google7 days agoView details
Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching servi…
marktechpost.com7 days agoView details
NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 delivered 58.5K successfully folded residues…
marktechpost.com7 days agoView details
OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call
OpenAI has released the Agents API in public beta. It gives developers the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. Developers run the agent’s compute in an OpenAI-managed sandbox, their own infrastructure, or a partner sandbox. Is it deployable? Yes. It is live for all d…
marktechpost.com7 days agoView details
Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B…
marktechpost.com8 days agoView details
Slack can now vibe-code interactive charts and reports inside chats
A new feature coming to Slack will allow you to build interactive reports, polls, dashboards, presentations, microsites, and other tools directly inside a chat. With Slackforce Surfaces, you can describe to Slackbot what you need, and it will use AI to gather information from relevant conversations and connected apps,…
theverge.com7 days agoView details
Is AI Actually Going to Kill Us All?
This week on “Uncanny Valley,” we dig into a former Anthropic researcher’s AI doomsday warning, the latest upgrades from Apple’s event, and the census report that claimed Trump won the 2020 election.
wired.com7 days agoView details
Schools are catching on to Big Tech’s playbook
It's the hot new thing in tech, and it's where all the jobs are. Students who don't learn to use it fall behind. And to help them catch up in time, its creators are graciously providing the resources and curriculum for learning it, often pro bono. That's the narrative AI companies are pitching schools on right now, bu…
theverge.com7 days agoView details
Universal Music is launching an AI music platform with ElevenLabs
Universal Music Group is launching a new AI-powered platform that will allow users to draw from its catalog of licensed music to create song remixes, mashups, and new takes on tracks, according to an announcement on Thursday. The record label is developing the platform through a multiyear licensing agreement with Elev…
theverge.com7 days agoView details
Meta’s Muse AI works and creeps me out
I turned my Muse assistant into a purple cat. | Screenshot: The Verge Meta has launched its new Muse assistant, marking the company's first real foray into AI-powered productivity tools. The company says its AI agent can "take the busywork off your plate" by helping you with online shopping, emails, trip-planning, and…
theverge.com7 days agoView details
Why the current tech backlash feels different
This interview has been lightly edited for length and clarity. Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. This is Nick Statt, senior producer. And I’m joined by our brand-new supervising producer, Greg Ott. Greg Ott: Good day, everyone. And Hi, Nilay. Nilay is here too.…
theverge.com7 days agoView details
Everything New You Can Do With Siri AI
When iOS 27 arrives, it will bring with it a fully revamped assistant for your iPhone.
wired.com7 days agoView details
Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online
InquiryIQ, a previously unreported prototype, tested a model from xAI, maker of Grok, to surface associates, social accounts, and other information about people identified through Clearview.
wired.com7 days agoView details
arXiv:2609.11020v1 Announce Type: new Abstract: We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixe…
arxiv.org7 days agoView details
A Fragility Spectrum for Recursive Language-Model Training
arXiv:2609.11149v1 Announce Type: new Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse.…
arxiv.org7 days agoView details
Domain-Specific Hallucination Detection in Large Language Models
arXiv:2609.11878v1 Announce Type: new Abstract: Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and tem…
arxiv.org7 days agoView details
arXiv:2609.11399v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevale…
arxiv.org7 days agoView details
SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations
arXiv:2609.11414v1 Announce Type: new Abstract: Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-turn dialogue, where routing per…
arxiv.org7 days agoView details
arXiv:2609.11244v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within…
arxiv.org7 days agoView details
Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu
arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used fo…
arxiv.org7 days agoView details
arXiv:2609.10810v1 Announce Type: new Abstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it…
arxiv.org7 days agoView details
arXiv:2609.10950v1 Announce Type: new Abstract: Recent multimodal sentiment analysis studies increasingly adopt text-centric fusion approaches to exploit the rich sentiment information inherent in the textual modality. However, these approaches often suffer from performance degradation during inference due to partiall…
arxiv.org7 days agoView details
arXiv:2609.10652v1 Announce Type: cross Abstract: Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is complex and requires trained exp…
arxiv.org7 days agoView details
LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection
arXiv:2609.10896v1 Announce Type: new Abstract: Speech-based automatic detection of Alzheimer's disease (AD) provides a non-invasive and scalable approach to early cognitive screening. AD affects both lexical-semantic organization and speech production, including atypical pauses and word elongations. However, existing…
arxiv.org7 days agoView details
ReGround: Grounding Reviewer Comments in Multimodal Evidence
arXiv:2609.11460v1 Announce Type: new Abstract: Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and largely focus on explicit, information-s…
arxiv.org7 days agoView details
arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were…
arxiv.org7 days agoView details
Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
arXiv:2609.11892v1 Announce Type: new Abstract: As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a co…
arxiv.org7 days agoView details
Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
arXiv:2609.11029v1 Announce Type: new Abstract: Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level text spans. To address this, we present an information-weighted cross-…
arxiv.org7 days agoView details
FlexComp: One Model for Every Ratio in Context Compression
arXiv:2609.11192v1 Announce Type: new Abstract: Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chose…
arxiv.org7 days agoView details
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
arXiv:2609.10745v1 Announce Type: new Abstract: Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph…
arxiv.org7 days agoView details
The widening evaluation gap in medical large language model research 2023 to 2026
arXiv:2609.11770v1 Announce Type: new Abstract: Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing…
arxiv.org7 days agoView details
The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge
arXiv:2609.11724v1 Announce Type: new Abstract: This paper details the Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages. Three approaches are explored. First, we fine-tune Voxtral-Mini-3B via…
arxiv.org7 days agoView details
Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System
arXiv:2609.10922v1 Announce Type: new Abstract: Auto-research agents have shown the potential to automate hypothesis generation, experiment execution, and iterative refinement. However, scaling this paradigm to industry-scale recommendation models introduces two challenges: (1) long feedback loops, where model trainin…
arxiv.org7 days agoView details
Structurally Speaking: Motif-Oriented Graph Captioning through Bidirectional Graph-Text Translation
arXiv:2609.10923v1 Announce Type: new Abstract: Graph captions should help readers understand graph structure, rather than simply translate adjacency matrices into long textual edge lists. A useful graph caption abstracts connectivity into recognizable motifs, such as hubs, paths, cycles, cliques, and bridges, because…
arxiv.org7 days agoView details
Distribution-aware Language Neuron Identification in Multilingual Large Language Models
arXiv:2609.10993v1 Announce Type: new Abstract: Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the entropy of each neuron's language-wise…
arxiv.org7 days agoView details
arXiv:2609.10996v1 Announce Type: new Abstract: Verbalized confidence, long dismissed as overconfident, coarse, and prone to round-number clustering, is now the more robust soft-scoring mechanism for LLM-as-a-Judge on top-tier proprietary models. Across SummEval, AggreFact, and HelpSteer2, spanning up to 18 LLMs, we s…
arxiv.org7 days agoView details
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation
arXiv:2609.11101v1 Announce Type: new Abstract: Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelity datasets, and coarse evaluati…
arxiv.org7 days agoView details
arXiv:2609.11117v1 Announce Type: new Abstract: Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment r…
arxiv.org7 days agoView details
Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment
arXiv:2609.10792v1 Announce Type: new Abstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accura…
arxiv.org7 days agoView details
Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures
arXiv:2609.10893v1 Announce Type: new Abstract: Recent advances in large language models have transformed human-computer interaction. Despite their fluency, these models often produce texts that are grammatically correct but semantically incoherent, containing contradictions or disruptions in logical flow. This work i…
arxiv.org7 days agoView details
SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs
arXiv:2609.10901v1 Announce Type: new Abstract: LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories…
arxiv.org7 days agoView details
Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking
arXiv:2609.10934v1 Announce Type: new Abstract: Turn-taking is a fundamental mechanism that governs when interlocutors speak and listen. Although Spoken Dialogue Systems (SDS) exploit a range of linguistic, acoustic, and non-verbal cues, they produce ill-timed responses in unscripted interaction. A central challenge i…
arxiv.org7 days agoView details
From Repetition to Recognition: Inductive Discovery of Disinformation Narratives
arXiv:2609.11128v1 Announce Type: new Abstract: In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as inductively inferring narrative labels from corpora, but its evaluation stays tied to predef…
arxiv.org7 days agoView details
Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting
arXiv:2609.11131v1 Announce Type: new Abstract: Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annota…
arxiv.org7 days agoView details
RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety
arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overal…
arxiv.org7 days agoView details
Distance generalization in transformers: why bother with positional encoding?
arXiv:2609.11913v1 Announce Type: new Abstract: Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training…
arxiv.org7 days agoView details
arXiv:2609.11545v1 Announce Type: new Abstract: Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, and content consistency using re…
arxiv.org7 days agoView details
Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
arXiv:2609.10702v1 Announce Type: new Abstract: Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million…
arxiv.org7 days agoView details
arXiv:2609.10715v1 Announce Type: new Abstract: We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, i…
arxiv.org7 days agoView details
CMNIE: An Information Extraction Benchmark for Chinese Military News
arXiv:2609.10722v1 Announce Type: new Abstract: Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event argument…
arxiv.org7 days agoView details
When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text
arXiv:2609.11067v1 Announce Type: new Abstract: Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain uncle…
arxiv.org7 days agoView details
Automated Identification of Competing Narratives in Political Discourse on Social Media
arXiv:2609.11202v1 Announce Type: new Abstract: Social media platforms have become central to shaping political discourse, serving as arenas where narratives form and evolve, influencing public opinion. Identifying and analyzing these narratives, particularly when they compete across different political ideologies, is…
arxiv.org7 days agoView details
SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ
arXiv:2609.11355v1 Announce Type: new Abstract: This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR i…
arxiv.org7 days agoView details
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization
arXiv:2609.11141v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate structured outputs, but their reliability remains unclear when those outputs must satisfy database-level constraints. We study this issue through database normalization, involving reasoning about functional d…
arxiv.org7 days agoView details
arXiv:2609.11246v1 Announce Type: new Abstract: Kurdish is spoken by millions of people, but little technology can read it aloud. A recent study released three Kurdish voices, 35 hours of recorded speech, and a paper describing the work, all free to download. This review checks how well those public files match the pa…
arxiv.org7 days agoView details
arXiv:2609.11247v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) remains constrained by modality imbalance, yet the field continues to rely on optimization-based balancing methods that promise more than they deliver. We provide three contributions: 1) a unified evaluation framework testing gradient…
arxiv.org7 days agoView details
arXiv:2609.11302v1 Announce Type: new Abstract: Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, where no prior benchmark for ALT…
arxiv.org7 days agoView details
MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions
arXiv:2609.11322v1 Announce Type: new Abstract: Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions…
arxiv.org7 days agoView details
arXiv:2609.11334v1 Announce Type: new Abstract: Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all ar…
arxiv.org7 days agoView details
On the Impact of Anonymization on the Performance of Large Language Models
arXiv:2609.11335v1 Announce Type: new Abstract: As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper pr…
arxiv.org7 days agoView details
Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study
arXiv:2609.11450v1 Announce Type: new Abstract: Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize it…
arxiv.org7 days agoView details
Structural priors for data-efficient language learning
arXiv:2609.11505v1 Announce Type: new Abstract: Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight i…
arxiv.org7 days agoView details
A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings
arXiv:2609.11620v1 Announce Type: new Abstract: High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across independently trained models. We…
arxiv.org7 days agoView details
Epistemic orientation predicts legislative effectiveness among members of the US Congress
arXiv:2609.11865v1 Announce Type: new Abstract: Truth and evidence-based communication provide important foundations for democratic governance, accountability, and collective decision-making. Prior work shows that evidence-oriented language in US congressional floor speeches has declined since the mid-1970s, alongside…
arxiv.org7 days agoView details
Structured Transforms for Low-Overhead Quantization of Language Models
arXiv:2609.11687v1 Announce Type: new Abstract: We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two c…
arxiv.org7 days agoView details
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that…
arxiv.org7 days agoView details
LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation
arXiv:2609.11739v1 Announce Type: new Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-ran…
arxiv.org7 days agoView details
Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs
arXiv:2609.11762v1 Announce Type: new Abstract: Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic enco…
arxiv.org7 days agoView details
Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing
arXiv:2609.11769v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can undo a known framing transform…
arxiv.org7 days agoView details
arXiv:2609.11786v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized. We present a switch a…
arxiv.org7 days agoView details
arXiv:2609.11838v1 Announce Type: new Abstract: Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the…
arxiv.org7 days agoView details
IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing
arXiv:2609.11851v1 Announce Type: new Abstract: Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Tra…
arxiv.org7 days agoView details
Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
arXiv:2609.11870v1 Announce Type: new Abstract: A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M wo…
arxiv.org7 days agoView details
Models
View all models →text-generation · transformers · safetensors · qwen3_5_text
huggingface.co8 days ago2169 ptsView details
text-generation · transformers · safetensors · qwen3_5_text
huggingface.co8 days ago767 ptsView details
image-text-to-text · safetensors · bailing_moe_v3_vl · image-text-to-text
huggingface.co7 days ago79 ptsView details
text-to-speech · sanotts · gguf · text-to-speech
huggingface.co8 days ago34 ptsView details
text-generation · transformers · safetensors · llama
huggingface.co8 days ago1 ptsView details
Ziyed-trading/marches-extracteur
gguf · endpoints_compatible · region:us
huggingface.co8 days agoView details
Open source
View all open source →<details open> vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle (#28618) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/46709648> **macOS/iOS:** - [macOS Apple Silicon (arm64)…
github.com7 days agoView details
## [3.13.0](https://github.com/openai/openai-python/compare/v3.12.0...v3.13.0) (2026-09-10) ### Features * **api:** add Agents API ([1c4284a](https://github.com/openai/openai-python/commit/1c4284a08294f734d57047585ab82e2e09d3a5bc))
github.com7 days agoView details
langchain-ai/langchain langchain-anthropic==1.7.2
Changes since langchain-anthropic==1.7.1 release(anthropic): 1.7.2 (#40387) fix(anthropic): preserve invalid tool use blocks (#40372)
github.com7 days agoView details
anthropics/anthropic-sdk-python v1.5.0
## 1.5.0 (2026-09-10) Full Changelog: [v1.4.0...v1.5.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.4.0...v1.5.0) ### Features * **api:** add auto mode tool permissions for Managed Agents ([62aa21b](https://github.com/anthropics/anthropic-sdk-python/commit/62aa…
github.com7 days agoView details
## [3.12.0](https://github.com/openai/openai-python/compare/v3.11.0...v3.12.0) (2026-09-10) ### Features * **api:** Add Live API ([0e4bfef](https://github.com/openai/openai-python/commit/0e4bfef9c79251fcf4926fd732129627bde050f1)) ### Bug Fixes * add aclose() to AsyncStream for s…
github.com7 days agoView details
<details open> ggml-cpu(s390x): add Q1_0 vector intrinsic support (#28606) * ggml-cpu: add `ggml_vec_dot_q1_0_q8_0` support Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ggml-cpu: clean up variable naming for understanding Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * docs:…
github.com8 days agoView details
<details open> model: fix all granite family parameter counts (#28643) * model: fix all granite family parameter counts Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * model: fix additional include, add missing `A` prefix for active experts Signed-off-by: Aaron Teo <aaron.teo1@i…
github.com8 days agoView details