Skip to content

Archive / 2026-09-10

September 10, 2026

  1. deepseek-ai/DeepSeek-V4.1-Flash

    image-text-to-text · transformers · safetensors · deepseek_v41

    huggingface.co8 days ago2911 ptsView details

  2. DeepSeek v4.1 Flash

    twitter.com8 days ago968 ptsView detailsJoin discussion

  3. OpenAI Agents API

    developers.openai.com7 days ago346 ptsView detailsJoin discussion

  4. Show HN: Bodily Oddities

    vester.si7 days ago335 ptsView detailsJoin discussion

  5. AI Is Breaking This Thing We Call Trust

    terriblesoftware.org7 days ago117 ptsView detailsJoin discussion

  6. Bending Spoons buying Miro for $1.355B

    investors.bendingspoons.com7 days ago103 ptsView detailsJoin discussion

  7. Mathematicians want proof OpenAI didn’t use their work

    Sam Altman, chief executive officer of OpenAI, during a media tour of the Stargate AI data center. | Bloomberg via Getty Images Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company's…

    theverge.com7 days ago77 ptsView details

  8. AI 2027 (2025)

    ai-2027.com7 days ago57 ptsView detailsJoin discussion

  9. Kagi Translate Is Back

    blog.kagi.com7 days ago43 ptsView detailsJoin discussion

  10. Stop externalizing the cost of your AI use to me

    thelastsoftwareengineer.substack.com7 days ago17 ptsView detailsJoin discussion

  11. Finding Slow Code with Wrapture

    grahamdumpleton.me8 days ago19 ptsView detailsJoin discussion

  12. The Age-Gating of History

    heatherburns.tech7 days ago17 ptsView detailsJoin discussion

  13. The Magic Behind Cubacadabra

    andrewarrow.dev7 days ago15 ptsView detailsJoin discussion

  14. OpenAI Wants to Know If an AI Industry Slowdown Would Even Be Legal

    AI leaders worry antitrust law could stand in the way of what they view as an increasingly urgent push to coordinate a slowdown in AI development.

    wired.com7 days ago13 ptsView detailsJoin discussion

  15. Primary source

    How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

    César de la Fuente’s lab uses Codex and ChatGPT to search living and extinct genomes for antimicrobial candidates to fight drug-resistant infections.

    openai.com7 days agoView details

  16. Primary source

    Now everyone can put data to work

    Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.

    openai.com7 days agoView details

  17. Primary source

    Introducing ChatGPT for Financial Services

    Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.

    openai.com8 days agoView details

  18. Primary source

    Expanding AI access and cyber defense for federal, state, local, and tribal governments

    OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.

    openai.com8 days agoView details

  19. Primary source

    ToolGrad: Efficient tool-use dataset generation with textual "gradients"

    Machine Intelligence

    research.google7 days agoView details

  20. Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

    Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching servi…

    marktechpost.com7 days agoView details

  21. NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

    NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 delivered 58.5K successfully folded residues…

    marktechpost.com7 days agoView details

  22. OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

    OpenAI has released the Agents API in public beta. It gives developers the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. Developers run the agent’s compute in an OpenAI-managed sandbox, their own infrastructure, or a partner sandbox. Is it deployable? Yes. It is live for all d…

    marktechpost.com7 days agoView details

  23. DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

    Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B…

    marktechpost.com8 days agoView details

  24. Slack can now vibe-code interactive charts and reports inside chats

    A new feature coming to Slack will allow you to build interactive reports, polls, dashboards, presentations, microsites, and other tools directly inside a chat. With Slackforce Surfaces, you can describe to Slackbot what you need, and it will use AI to gather information from relevant conversations and connected apps,…

    theverge.com7 days agoView details

  25. Is AI Actually Going to Kill Us All?

    This week on “Uncanny Valley,” we dig into a former Anthropic researcher’s AI doomsday warning, the latest upgrades from Apple’s event, and the census report that claimed Trump won the 2020 election.

    wired.com7 days agoView details

  26. Schools are catching on to Big Tech’s playbook

    It's the hot new thing in tech, and it's where all the jobs are. Students who don't learn to use it fall behind. And to help them catch up in time, its creators are graciously providing the resources and curriculum for learning it, often pro bono. That's the narrative AI companies are pitching schools on right now, bu…

    theverge.com7 days agoView details

  27. Universal Music is launching an AI music platform with ElevenLabs

    Universal Music Group is launching a new AI-powered platform that will allow users to draw from its catalog of licensed music to create song remixes, mashups, and new takes on tracks, according to an announcement on Thursday. The record label is developing the platform through a multiyear licensing agreement with Elev…

    theverge.com7 days agoView details

  28. Meta’s Muse AI works and creeps me out

    I turned my Muse assistant into a purple cat. | Screenshot: The Verge Meta has launched its new Muse assistant, marking the company's first real foray into AI-powered productivity tools. The company says its AI agent can "take the busywork off your plate" by helping you with online shopping, emails, trip-planning, and…

    theverge.com7 days agoView details

  29. Why the current tech backlash feels different

    This interview has been lightly edited for length and clarity. Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. This is Nick Statt, senior producer. And I’m joined by our brand-new supervising producer, Greg Ott. Greg Ott: Good day, everyone. And Hi, Nilay. Nilay is here too.…

    theverge.com7 days agoView details

  30. Everything New You Can Do With Siri AI

    When iOS 27 arrives, it will bring with it a fully revamped assistant for your iPhone.

    wired.com7 days agoView details

  31. Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online

    InquiryIQ, a previously unreported prototype, tested a model from xAI, maker of Grok, to surface associates, social accounts, and other information about people identified through Clearview.

    wired.com7 days agoView details

  32. K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models

    arXiv:2609.11020v1 Announce Type: new Abstract: We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixe…

    arxiv.org7 days agoView details

  33. A Fragility Spectrum for Recursive Language-Model Training

    arXiv:2609.11149v1 Announce Type: new Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse.…

    arxiv.org7 days agoView details

  34. Domain-Specific Hallucination Detection in Large Language Models

    arXiv:2609.11878v1 Announce Type: new Abstract: Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and tem…

    arxiv.org7 days agoView details

  35. TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

    arXiv:2609.11399v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevale…

    arxiv.org7 days agoView details

  36. SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

    arXiv:2609.11414v1 Announce Type: new Abstract: Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-turn dialogue, where routing per…

    arxiv.org7 days agoView details

  37. OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

    arXiv:2609.11244v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within…

    arxiv.org7 days agoView details

  38. Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

    arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used fo…

    arxiv.org7 days agoView details

  39. Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

    arXiv:2609.10810v1 Announce Type: new Abstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it…

    arxiv.org7 days agoView details

  40. Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction

    arXiv:2609.10950v1 Announce Type: new Abstract: Recent multimodal sentiment analysis studies increasingly adopt text-centric fusion approaches to exploit the rich sentiment information inherent in the textual modality. However, these approaches often suffer from performance degradation during inference due to partiall…

    arxiv.org7 days agoView details

  41. Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

    arXiv:2609.10652v1 Announce Type: cross Abstract: Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is complex and requires trained exp…

    arxiv.org7 days agoView details

  42. LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

    arXiv:2609.10896v1 Announce Type: new Abstract: Speech-based automatic detection of Alzheimer's disease (AD) provides a non-invasive and scalable approach to early cognitive screening. AD affects both lexical-semantic organization and speech production, including atypical pauses and word elongations. However, existing…

    arxiv.org7 days agoView details

  43. ReGround: Grounding Reviewer Comments in Multimodal Evidence

    arXiv:2609.11460v1 Announce Type: new Abstract: Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and largely focus on explicit, information-s…

    arxiv.org7 days agoView details

  44. Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

    arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were…

    arxiv.org7 days agoView details

  45. Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

    arXiv:2609.11892v1 Announce Type: new Abstract: As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a co…

    arxiv.org7 days agoView details

  46. Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss

    arXiv:2609.11029v1 Announce Type: new Abstract: Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level text spans. To address this, we present an information-weighted cross-…

    arxiv.org7 days agoView details

  47. FlexComp: One Model for Every Ratio in Context Compression

    arXiv:2609.11192v1 Announce Type: new Abstract: Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chose…

    arxiv.org7 days agoView details

  48. Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

    arXiv:2609.10745v1 Announce Type: new Abstract: Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph…

    arxiv.org7 days agoView details

  49. The widening evaluation gap in medical large language model research 2023 to 2026

    arXiv:2609.11770v1 Announce Type: new Abstract: Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing…

    arxiv.org7 days agoView details

  50. The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

    arXiv:2609.11724v1 Announce Type: new Abstract: This paper details the Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages. Three approaches are explored. First, we fine-tune Voxtral-Mini-3B via…

    arxiv.org7 days agoView details

  51. Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System

    arXiv:2609.10922v1 Announce Type: new Abstract: Auto-research agents have shown the potential to automate hypothesis generation, experiment execution, and iterative refinement. However, scaling this paradigm to industry-scale recommendation models introduces two challenges: (1) long feedback loops, where model trainin…

    arxiv.org7 days agoView details

  52. Structurally Speaking: Motif-Oriented Graph Captioning through Bidirectional Graph-Text Translation

    arXiv:2609.10923v1 Announce Type: new Abstract: Graph captions should help readers understand graph structure, rather than simply translate adjacency matrices into long textual edge lists. A useful graph caption abstracts connectivity into recognizable motifs, such as hubs, paths, cycles, cliques, and bridges, because…

    arxiv.org7 days agoView details

  53. Distribution-aware Language Neuron Identification in Multilingual Large Language Models

    arXiv:2609.10993v1 Announce Type: new Abstract: Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the entropy of each neuron's language-wise…

    arxiv.org7 days agoView details

  54. Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

    arXiv:2609.10996v1 Announce Type: new Abstract: Verbalized confidence, long dismissed as overconfident, coarse, and prone to round-number clustering, is now the more robust soft-scoring mechanism for LLM-as-a-Judge on top-tier proprietary models. Across SummEval, AggreFact, and HelpSteer2, spanning up to 18 LLMs, we s…

    arxiv.org7 days agoView details

  55. ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation

    arXiv:2609.11101v1 Announce Type: new Abstract: Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelity datasets, and coarse evaluati…

    arxiv.org7 days agoView details

  56. Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

    arXiv:2609.11117v1 Announce Type: new Abstract: Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment r…

    arxiv.org7 days agoView details

  57. Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

    arXiv:2609.10792v1 Announce Type: new Abstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accura…

    arxiv.org7 days agoView details

  58. Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures

    arXiv:2609.10893v1 Announce Type: new Abstract: Recent advances in large language models have transformed human-computer interaction. Despite their fluency, these models often produce texts that are grammatically correct but semantically incoherent, containing contradictions or disruptions in logical flow. This work i…

    arxiv.org7 days agoView details

  59. SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs

    arXiv:2609.10901v1 Announce Type: new Abstract: LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories…

    arxiv.org7 days agoView details

  60. Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking

    arXiv:2609.10934v1 Announce Type: new Abstract: Turn-taking is a fundamental mechanism that governs when interlocutors speak and listen. Although Spoken Dialogue Systems (SDS) exploit a range of linguistic, acoustic, and non-verbal cues, they produce ill-timed responses in unscripted interaction. A central challenge i…

    arxiv.org7 days agoView details

  61. From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

    arXiv:2609.11128v1 Announce Type: new Abstract: In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as inductively inferring narrative labels from corpora, but its evaluation stays tied to predef…

    arxiv.org7 days agoView details

  62. Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

    arXiv:2609.11131v1 Announce Type: new Abstract: Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annota…

    arxiv.org7 days agoView details

  63. RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

    arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overal…

    arxiv.org7 days agoView details

  64. Distance generalization in transformers: why bother with positional encoding?

    arXiv:2609.11913v1 Announce Type: new Abstract: Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training…

    arxiv.org7 days agoView details

  65. Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

    arXiv:2609.11545v1 Announce Type: new Abstract: Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, and content consistency using re…

    arxiv.org7 days agoView details

  66. Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

    arXiv:2609.10702v1 Announce Type: new Abstract: Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million…

    arxiv.org7 days agoView details

  67. NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

    arXiv:2609.10715v1 Announce Type: new Abstract: We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, i…

    arxiv.org7 days agoView details

  68. CMNIE: An Information Extraction Benchmark for Chinese Military News

    arXiv:2609.10722v1 Announce Type: new Abstract: Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event argument…

    arxiv.org7 days agoView details

  69. When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

    arXiv:2609.11067v1 Announce Type: new Abstract: Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain uncle…

    arxiv.org7 days agoView details

  70. Automated Identification of Competing Narratives in Political Discourse on Social Media

    arXiv:2609.11202v1 Announce Type: new Abstract: Social media platforms have become central to shaping political discourse, serving as arenas where narratives form and evolve, influencing public opinion. Identifying and analyzing these narratives, particularly when they compete across different political ideologies, is…

    arxiv.org7 days agoView details

  71. SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

    arXiv:2609.11355v1 Announce Type: new Abstract: This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR i…

    arxiv.org7 days agoView details

  72. Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization

    arXiv:2609.11141v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate structured outputs, but their reliability remains unclear when those outputs must satisfy database-level constraints. We study this issue through database normalization, involving reasoning about functional d…

    arxiv.org7 days agoView details

  73. Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study

    arXiv:2609.11246v1 Announce Type: new Abstract: Kurdish is spoken by millions of people, but little technology can read it aloud. A recent study released three Kurdish voices, 35 hours of recorded speech, and a paper describing the work, all free to download. This review checks how well those public files match the pa…

    arxiv.org7 days agoView details

  74. The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods

    arXiv:2609.11247v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) remains constrained by modality imbalance, yet the field continues to rely on optimization-based balancing methods that promise more than they deliver. We provide three contributions: 1) a unified evaluation framework testing gradient…

    arxiv.org7 days agoView details

  75. Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

    arXiv:2609.11302v1 Announce Type: new Abstract: Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, where no prior benchmark for ALT…

    arxiv.org7 days agoView details

  76. MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

    arXiv:2609.11322v1 Announce Type: new Abstract: Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions…

    arxiv.org7 days agoView details

  77. E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

    arXiv:2609.11334v1 Announce Type: new Abstract: Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all ar…

    arxiv.org7 days agoView details

  78. On the Impact of Anonymization on the Performance of Large Language Models

    arXiv:2609.11335v1 Announce Type: new Abstract: As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper pr…

    arxiv.org7 days agoView details

  79. Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study

    arXiv:2609.11450v1 Announce Type: new Abstract: Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize it…

    arxiv.org7 days agoView details

  80. Structural priors for data-efficient language learning

    arXiv:2609.11505v1 Announce Type: new Abstract: Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight i…

    arxiv.org7 days agoView details

  81. A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

    arXiv:2609.11620v1 Announce Type: new Abstract: High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across independently trained models. We…

    arxiv.org7 days agoView details

  82. Epistemic orientation predicts legislative effectiveness among members of the US Congress

    arXiv:2609.11865v1 Announce Type: new Abstract: Truth and evidence-based communication provide important foundations for democratic governance, accountability, and collective decision-making. Prior work shows that evidence-oriented language in US congressional floor speeches has declined since the mid-1970s, alongside…

    arxiv.org7 days agoView details

  83. Structured Transforms for Low-Overhead Quantization of Language Models

    arXiv:2609.11687v1 Announce Type: new Abstract: We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two c…

    arxiv.org7 days agoView details

  84. Negative Self-Distillation: Learning to Reason by Avoiding Flaws

    arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that…

    arxiv.org7 days agoView details

  85. LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

    arXiv:2609.11739v1 Announce Type: new Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-ran…

    arxiv.org7 days agoView details

  86. Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

    arXiv:2609.11762v1 Announce Type: new Abstract: Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic enco…

    arxiv.org7 days agoView details

  87. Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing

    arXiv:2609.11769v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can undo a known framing transform…

    arxiv.org7 days agoView details

  88. Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech

    arXiv:2609.11786v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized. We present a switch a…

    arxiv.org7 days agoView details

  89. Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

    arXiv:2609.11838v1 Announce Type: new Abstract: Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the…

    arxiv.org7 days agoView details

  90. IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

    arXiv:2609.11851v1 Announce Type: new Abstract: Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Tra…

    arxiv.org7 days agoView details

  91. Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

    arXiv:2609.11870v1 Announce Type: new Abstract: A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M wo…

    arxiv.org7 days agoView details

  1. TokenRhythm/NeoHorse-1-4B

    text-generation · transformers · safetensors · qwen3_5_text

    huggingface.co8 days ago2169 ptsView details

  2. TokenRhythm/NeoHorse-1-9B

    text-generation · transformers · safetensors · qwen3_5_text

    huggingface.co8 days ago767 ptsView details

  3. inclusionAI/Ling-3.0-flash-VL

    image-text-to-text · safetensors · bailing_moe_v3_vl · image-text-to-text

    huggingface.co7 days ago79 ptsView details

  4. ampixa/sanoTTS

    text-to-speech · sanotts · gguf · text-to-speech

    huggingface.co8 days ago34 ptsView details

  5. kikusuka/breezy-78m-pretrain

    text-generation · transformers · safetensors · llama

    huggingface.co8 days ago1 ptsView details

  6. Ziyed-trading/marches-extracteur

    gguf · endpoints_compatible · region:us

    huggingface.co8 days agoView details

  1. ggml-org/llama.cpp b10901

    <details open> vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle (#28618) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/46709648> **macOS/iOS:** - [macOS Apple Silicon (arm64)…

    github.com7 days agoView details

  2. openai/openai-python v3.13.0

    ## [3.13.0](https://github.com/openai/openai-python/compare/v3.12.0...v3.13.0) (2026-09-10) ### Features * **api:** add Agents API ([1c4284a](https://github.com/openai/openai-python/commit/1c4284a08294f734d57047585ab82e2e09d3a5bc))

    github.com7 days agoView details

  3. langchain-ai/langchain langchain-anthropic==1.7.2

    Changes since langchain-anthropic==1.7.1 release(anthropic): 1.7.2 (#40387) fix(anthropic): preserve invalid tool use blocks (#40372)

    github.com7 days agoView details

  4. anthropics/anthropic-sdk-python v1.5.0

    ## 1.5.0 (2026-09-10) Full Changelog: [v1.4.0...v1.5.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.4.0...v1.5.0) ### Features * **api:** add auto mode tool permissions for Managed Agents ([62aa21b](https://github.com/anthropics/anthropic-sdk-python/commit/62aa…

    github.com7 days agoView details

  5. openai/openai-python v3.12.0

    ## [3.12.0](https://github.com/openai/openai-python/compare/v3.11.0...v3.12.0) (2026-09-10) ### Features * **api:** Add Live API ([0e4bfef](https://github.com/openai/openai-python/commit/0e4bfef9c79251fcf4926fd732129627bde050f1)) ### Bug Fixes * add aclose() to AsyncStream for s…

    github.com7 days agoView details

  6. ggml-org/llama.cpp b10886

    <details open> ggml-cpu(s390x): add Q1_0 vector intrinsic support (#28606) * ggml-cpu: add `ggml_vec_dot_q1_0_q8_0` support Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ggml-cpu: clean up variable naming for understanding Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * docs:…

    github.com8 days agoView details

  7. ggml-org/llama.cpp b10885

    <details open> model: fix all granite family parameter counts (#28643) * model: fix all granite family parameter counts Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * model: fix additional include, add missing `A` prefix for active experts Signed-off-by: Aaron Teo <aaron.teo1@i…

    github.com8 days agoView details