Skip to content

Archive / 2026-08-20

August 20, 2026

  1. Don't paste the AI, please

    dontpastetheai.com29 days ago1038 ptsView detailsJoin discussion

  2. Consumer Rights Wiki

    consumerrights.wiki28 days ago293 ptsView detailsJoin discussion

  3. Why aren't smart people happier? (2022)

    experimental-history.com28 days ago266 ptsView detailsJoin discussion

  4. Ox Alpha

    openrouter.ai28 days ago250 ptsView detailsJoin discussion

  5. Stop eating Lady Gaga's Oreos

    experimental-history.com28 days ago202 ptsView detailsJoin discussion

  6. An elliptic curve of rank ≥ 30

    elliptic-rank.icarm.cloud28 days ago94 ptsView detailsJoin discussion

  7. Artificial Intelligence Policy

    law.berkeley.edu28 days ago42 ptsView detailsJoin discussion

  8. Primary source

    Introducing Intelligence Age

    Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

    openai.com29 days ago32 ptsView details

  9. AI at Home Part 2: Multi-GPU Drifting

    jdagostino.github.io28 days ago29 ptsView detailsJoin discussion

  10. Could AIs Become Conscious?

    economist.com28 days ago26 ptsView detailsJoin discussion

  11. Primary source

    Introducing Intelligence Age

    Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

    openai.com29 days agoView details

  12. Primary source

    Up to 3.2x Faster Inference with LFM2.5-DSpark

    huggingface.co28 days agoView details

  13. Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text

    S1-mini is a 462 MB open-weights normalizer that sits after ASR, removing fillers and resolving self-corrections locally. The post Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text appeared first on MarkTechPost.

    marktechpost.com28 days agoView details

  14. Meet UPDF: A Lightweight Adobe Alternative Built for the Agentic Era

    PDFs are easy to read and hard to change. AI can now summarize a 90-page contract in seconds, but it still won't rewrite the source file cleanly. UPDF is built for that second half: direct editing, 14-format conversion, 38-language OCR, and ten AI agents shipped in version 2.5. The post Meet UPDF: A Lightweight Adobe…

    marktechpost.com28 days agoView details

  15. Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs

    Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output. The post Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs appeared first on MarkTechPost.

    marktechpost.com28 days agoView details

  16. Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA

    This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure g…

    marktechpost.com29 days agoView details

  17. Google Discover is getting an AI chatbot-tuned feed

    Google will soon allow you to customize your Discover feed by describing what you want to see. The new feature, rolling out to the Google app in the "coming days," will use AI to automatically tweak your feed and "remember" your preferences for future visits. You'll find the option within the three-dot menu on your Di…

    theverge.com28 days agoView details

  18. Silicon Valley Doesn't Get Why You Hate AI

    Technology leaders don’t seem to understand society’s gripes about AI, but boy, are they posting through it.

    wired.com28 days agoView details

  19. It’s Greg Brockman’s OpenAI now

    OpenAI has had a hell of a year. The company spent months battling former cofounder Elon Musk in a sensational jury trial, was hit with a high-profile trade secrets lawsuit from Apple, and faced widespread scrutiny after an unreleased model hacked another AI company. As it prepares for an IPO, a steady string of execu…

    theverge.com28 days agoView details

  20. Welcome to the AI crisis in math

    Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it. OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in t…

    theverge.com28 days agoView details

  21. Slack is launching collaborative vibe-coding channels

    Slack is introducing dedicated channels where teams can vibe-code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includes open, project-specific code channels with dedicated user tabs, alongside features that compare coding changes and preview HTML output be…

    theverge.com28 days agoView details

  22. ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

    arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed…

    arxiv.org28 days agoView details

  23. A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment

    arXiv:2608.19199v1 Announce Type: new Abstract: We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge for members of a Community of Practice (CoP) related to the Oil and Gas sector. An evaluation of a first prototype involving 75 pro…

    arxiv.org28 days agoView details

  24. Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life

    arXiv:2608.19218v1 Announce Type: new Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). In this paper, we investigate…

    arxiv.org28 days agoView details

  25. Can Agent Memory Systems Track Evolving State?

    arXiv:2608.19652v1 Announce Type: cross Abstract: As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the worl…

    arxiv.org28 days agoView details

  26. Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa

    arXiv:2608.19200v1 Announce Type: new Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information. Automatic text summarization (ATS), driven by advancements in natural language processing (NLP), has developed rapidly in recent years. ATS methods…

    arxiv.org28 days agoView details

  27. Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses

    arXiv:2608.19206v1 Announce Type: new Abstract: Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development…

    arxiv.org28 days agoView details

  28. When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

    arXiv:2608.20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we…

    arxiv.org28 days agoView details

  29. Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping

    arXiv:2608.19220v1 Announce Type: new Abstract: Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-grou…

    arxiv.org28 days agoView details

  30. Reliable Financial Named Entity Recognition under Domain Shift

    arXiv:2608.19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate when the input distribution changes. We st…

    arxiv.org28 days agoView details

  31. PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

    arXiv:2608.19598v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitatio…

    arxiv.org28 days agoView details

  32. SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

    arXiv:2608.19799v1 Announce Type: new Abstract: Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely em…

    arxiv.org28 days agoView details

  33. The Asymmetric Harms of LLM Compression

    arXiv:2608.19670v1 Announce Type: new Abstract: Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11 compression methods to investigate the effect…

    arxiv.org28 days agoView details

  34. Stopping and Routing LLM Judge Panels

    arXiv:2608.19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is best, but which judges should be called, on…

    arxiv.org28 days agoView details

  35. SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit

    arXiv:2608.19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing. Diachronic corpus research instead examines interpretable dimensions such as syntactic be…

    arxiv.org28 days agoView details

  36. Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder

    arXiv:2608.19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models have been virtually non-existent. We prese…

    arxiv.org28 days agoView details

  37. LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment

    arXiv:2608.19800v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap b…

    arxiv.org28 days agoView details

  38. Automatic bioinformatic software named entity recognition from literature

    arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of…

    arxiv.org28 days agoView details

  39. Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

    arXiv:2608.19203v1 Announce Type: new Abstract: Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. Some heads may rely mainly on nearby lexical or syntactic context, while others may depend on longer-range relations such as entit…

    arxiv.org28 days agoView details

  40. Inducing Task Models from Computer-Use Traces

    arXiv:2608.20319v1 Announce Type: new Abstract: Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as computer-use agents enter real work, where ag…

    arxiv.org28 days agoView details

  41. Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

    arXiv:2608.19564v1 Announce Type: new Abstract: Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used only in the current context, re-verified, o…

    arxiv.org28 days agoView details

  42. Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

    arXiv:2608.19611v1 Announce Type: new Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a…

    arxiv.org28 days agoView details

  43. Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories

    arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis show…

    arxiv.org28 days agoView details

  44. Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

    arXiv:2608.19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-orie…

    arxiv.org28 days agoView details

  45. Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

    arXiv:2608.19207v1 Announce Type: new Abstract: Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn, leaving system-message adherence in multim…

    arxiv.org28 days agoView details

  46. When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

    arXiv:2608.19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled in…

    arxiv.org28 days agoView details

  47. StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

    arXiv:2608.19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commenta…

    arxiv.org28 days agoView details

  48. Active Inference as Context Acquisition for AI Agents

    arXiv:2608.19202v1 Announce Type: cross Abstract: Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying question, retrieval call, tool call, or p…

    arxiv.org28 days agoView details

  49. SABET-QA: Temporal Knowledge Graph Question Answering

    arXiv:2608.20083v1 Announce Type: new Abstract: Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose SABET-QA, a framework that iteratively refines…

    arxiv.org28 days agoView details

  50. Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems

    arXiv:2608.19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort an…

    arxiv.org28 days agoView details

  51. Does Listening Matter? Backchanneling and Nodding in AI Clone

    arXiv:2608.19527v1 Announce Type: cross Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence and authenticity. We integrated verbal backchann…

    arxiv.org28 days agoView details

  52. A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

    arXiv:2608.19875v1 Announce Type: new Abstract: Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on…

    arxiv.org28 days agoView details

  53. A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

    arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models an…

    arxiv.org28 days agoView details

  54. Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving Transformations

    arXiv:2608.19369v1 Announce Type: new Abstract: Statistical watermarks for language models live in the freedom of the signifier: they choose among tokens that are nearly equivalent in meaning, and they are therefore eroded by exactly those transformations which move the form of a text while leaving its content in plac…

    arxiv.org28 days agoView details

  55. ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents

    arXiv:2608.19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReCache}, a framework for independently ca…

    arxiv.org28 days agoView details

  56. Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection

    arXiv:2608.19942v1 Announce Type: new Abstract: Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challe…

    arxiv.org28 days agoView details

  57. Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models

    arXiv:2608.19211v1 Announce Type: new Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but al…

    arxiv.org28 days agoView details

  58. TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling

    arXiv:2608.19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly…

    arxiv.org28 days agoView details

  59. Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

    arXiv:2608.19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information int…

    arxiv.org28 days agoView details

  60. FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

    arXiv:2608.20153v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on fron…

    arxiv.org28 days agoView details

  61. Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

    arXiv:2608.20169v1 Announce Type: new Abstract: We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the und…

    arxiv.org28 days agoView details

  62. Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

    arXiv:2608.20281v1 Announce Type: new Abstract: Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge…

    arxiv.org28 days agoView details

  63. G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

    arXiv:2608.20331v1 Announce Type: new Abstract: Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do…

    arxiv.org28 days agoView details

  64. Outcome Monitors: Recovery Affordances for Silent Tool Failures

    arXiv:2608.19303v1 Announce Type: cross Abstract: When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect violations of outcome contracts mined fr…

    arxiv.org28 days agoView details

  65. HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization

    arXiv:2608.19430v1 Announce Type: cross Abstract: Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios. Existing scoring systems and ranking methods typically assume a fixed criterion. In practice…

    arxiv.org28 days agoView details

  66. Are LLMs becoming similarly creative? Evidence from three years of models

    arXiv:2608.19437v1 Announce Type: new Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly sup…

    arxiv.org28 days agoView details

  67. DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule

    arXiv:2608.19491v1 Announce Type: cross Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequentl…

    arxiv.org28 days agoView details

  68. Measuring What a Specification Determines: A Formal Semantic-Block Model and an Execution-Judged Benchmark

    arXiv:2608.19475v1 Announce Type: cross Abstract: This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification quality independently of model capability. A specification is represented as a structure comprising semantic blocks, dependency relation…

    arxiv.org28 days agoView details

  69. From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

    arXiv:2608.19535v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Con…

    arxiv.org28 days agoView details

  70. Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)

    arXiv:2608.19526v1 Announce Type: new Abstract: Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect investment decisions. This projec…

    arxiv.org28 days agoView details

  71. When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models

    arXiv:2608.19529v1 Announce Type: new Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure, they lie outside the linguistic token sp…

    arxiv.org28 days agoView details

  72. Projector Is All You Train

    arXiv:2608.19726v1 Announce Type: new Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it…

    arxiv.org28 days agoView details

  73. PersonalBench: Measuring the Authorship Gap in LLM Personalization

    arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing. We introduce PersonalBench,…

    arxiv.org28 days agoView details

  74. FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

    arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through inst…

    arxiv.org28 days agoView details

  75. Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

    arXiv:2608.19893v1 Announce Type: new Abstract: Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation:…

    arxiv.org28 days agoView details

  76. Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

    arXiv:2608.19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It…

    arxiv.org28 days agoView details

  77. One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

    arXiv:2608.19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agen…

    arxiv.org28 days agoView details

  78. Auditing Cross-Lingual Fairness in Language Model Watermarking

    arXiv:2608.20047v1 Announce Type: new Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on E…

    arxiv.org28 days agoView details

  79. NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection

    arXiv:2608.19212v1 Announce Type: new Abstract: Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC mi…

    arxiv.org28 days agoView details

  80. HealMed: Multilingual Evaluation of Large Language Models in Medicine

    arXiv:2608.19981v1 Announce Type: new Abstract: We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchm…

    arxiv.org28 days agoView details

  81. OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

    arXiv:2608.20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built from 38,104 atomic, source-anchored fac…

    arxiv.org28 days agoView details

  1. tencent/UI-Mate-27B

    image-text-to-text · transformers · safetensors · qwen3_5

    huggingface.co28 days ago79 ptsView details

  2. lightx2v/Minimax-h3-Turbo-SLA

    image-to-video · video-generation · image-to-video · audio-video-generation

    huggingface.co29 days ago78 ptsView details

  3. orcarouter/Qwen3.8-27B-Uncensored

    image-text-to-text · transformers · safetensors · qwen3_5

    huggingface.co29 days ago48 ptsView details

  4. ReliquaryForge/qwen3-4b-base-dapo-v4

    text-generation · transformers · safetensors · qwen3

    huggingface.co29 days ago7 ptsView details

  5. skypro1111/gemma-3-270m-uk-verbalizer

    text-generation · transformers · onnx · safetensors

    huggingface.co29 days ago2 ptsView details

  6. software-mansion/react-native-executorch-efficientnet-v2-s

    executorch · license:other · region:us

    huggingface.co29 days ago1 ptsView details

  7. orpe42/deberta_MP_dynamic

    text-classification · transformers · tensorboard · safetensors

    huggingface.co29 days agoView details

  1. ggml-org/llama.cpp b10545

    <details open> metal : clamp K extent in tensor API mat-mat kernel for K not a multiple of 32 (#27450) The Tensor API mat-mat path of kernel_mul_mm (GGML_METAL_HAS_TENSOR) fed a static K=32 tile to the matmul2d op on every iteration. On the last, partial K tile (ne00 % 32 != 0)…

    github.com28 days agoView details

  2. ggml-org/llama.cpp b10541

    <details open> mtmd: add --mmproj-device argument (#23255) * feat: add --mmproj-device arg & backwards compatible MTMD_BACKEND_DEVICE env var * feat: load mmproj device backend immediately, add -mmdev shortflag * fix: its a pointer now get the name * clean up * gen docs * nits -…

    github.com28 days agoView details

  3. ggml-org/llama.cpp b10539

    <details open> vulkan: FA MMQ should use fp32 for Q quantization calculations (#27413) Codex found that qd could be a denorm and 1/qd would overflow. </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/42034841> *…

    github.com28 days agoView details

  4. anthropics/anthropic-sdk-python v1.0.0

    ## 1.0.0 (2026-08-20) Full Changelog: [v0.125.0...v1.0.0](https://github.com/anthropics/anthropic-sdk-python/compare/v0.125.0...v1.0.0) ### ⚠ BREAKING CHANGES * **client:** upgrade to httpx2 and some minor breaking changes. See MIGRATION.md for details ### Features * **client:**…

    github.com28 days agoView details

  5. langchain-ai/langchain langchain-fireworks==1.6.0

    Changes since langchain-fireworks==1.5.2 release(fireworks): 1.6.0 (#39810) feat(fireworks): add document reranking (#39732) fix(fireworks): filter invalid tool calls from v1 content (#39805) feat(core): add standard model exception types (#39538) chore(model-profiles): refresh…

    github.com28 days agoView details

  6. langchain-ai/langchain langchain==1.3.16

    Changes since langchain==1.3.15 release(langchain): 1.3.16 (#39806) feat(core): add standard model exception types (#39538) feat(langchain): support custom token_counter in ContextEditingMiddleware (#39754) fix(langchain): re-raise non-retryable exceptions in ModelRetryMiddlewar…

    github.com28 days agoView details

  7. langchain-ai/langchain langchain-anthropic==1.6.1

    Changes since langchain-anthropic==1.6.0 release(anthropic): 1.6.1 (#39804) fix(anthropic): filter invalid tool calls from v1 content (#39803)

    github.com28 days agoView details

  8. ggml-org/llama.cpp b10509

    <details open> ggml: add ggml_rope_set_offset (+ metal support) (#27120) * add params * cpu kernel * metal kernel * add test backend ops * gate other backends * ggml: (cuda) support ggml_rope_set_offset (#27121) * rm cuda supports_op guard, fix webgpu clang-format * ggml: suppor…

    github.com29 days agoView details