Archive / 2026-08-20
August 20, 2026
News
View all news →Aaron Swartz was prosecuted for scraping, while Meta does it without consequence
blog.curiousquail.com28 days ago1671 ptsView detailsJoin discussion
dontpastetheai.com29 days ago1038 ptsView detailsJoin discussion
AI companies destroy physical books – let's scan rare books before it's too late
annas-archive.gl28 days ago623 ptsView detailsJoin discussion
Show HN: Huzzah – a novel approach to coding with AI
danielvaughn.dev28 days ago384 ptsView detailsJoin discussion
Watching TikTok and Instagram deactivates the cognitive control network: Study
rathbiotaclan.com28 days ago354 ptsView detailsJoin discussion
Vomit: Clean up Claude 5's token output with a separate LLM
github.com28 days ago305 ptsView detailsJoin discussion
consumerrights.wiki28 days ago293 ptsView detailsJoin discussion
Why aren't smart people happier? (2022)
experimental-history.com28 days ago266 ptsView detailsJoin discussion
openrouter.ai28 days ago250 ptsView detailsJoin discussion
Anti-AI fonts are useless and harmful
blog.yaros.ae28 days ago210 ptsView detailsJoin discussion
experimental-history.com28 days ago202 ptsView detailsJoin discussion
Copyright does not protect AI-generated content in EU
mathstodon.xyz28 days ago189 ptsView detailsJoin discussion
Micron announces $10B research hub in Boise
investors.micron.com28 days ago129 ptsView detailsJoin discussion
Autolith: A programming agent with a live runtime
lambda-symbolics.com28 days ago126 ptsView detailsJoin discussion
An elliptic curve of rank ≥ 30
elliptic-rank.icarm.cloud28 days ago94 ptsView detailsJoin discussion
AI didn't erase the junior engineer's value, it increased it it
franciscotrindade.me28 days ago89 ptsView detailsJoin discussion
Show HN: Open-source Stripe Connect alternative
zoneless.com28 days ago83 ptsView detailsJoin discussion
Show HN: Check if any of the $656M in unclaimed royalties at The MLC is yours
pub.doub.ly28 days ago73 ptsView detailsJoin discussion
Guess which of these LLM outputs is watermarked
sgoedecke.github.io28 days ago71 ptsView detailsJoin discussion
Launch HN: Vendo (YC S26) – Let users build features on top of your product
github.com28 days ago61 ptsView detailsJoin discussion
Show HN: We chased a weather balloon across Montana and never found it
radi8.dev28 days ago60 ptsView detailsJoin discussion
Show HN: Omacosy – Omarchy-style tiling desktop for macOS, no SIP
github.com28 days ago53 ptsView detailsJoin discussion
Artificial Intelligence Policy
law.berkeley.edu28 days ago42 ptsView detailsJoin discussion
- Primary source
Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.
openai.com29 days ago32 ptsView details
AI at Home Part 2: Multi-GPU Drifting
jdagostino.github.io28 days ago29 ptsView detailsJoin discussion
ProgramBench Vetted: Reverse Engineering from a Runnable Binary
vetto.ai28 days ago27 ptsView detailsJoin discussion
economist.com28 days ago26 ptsView detailsJoin discussion
The city of Cupertino is 70.9% Asian. A new TV show set there ignores this
sfgate.com28 days ago21 ptsView detailsJoin discussion
Protesters haul a guillotine to city council meeting about an AI data center
tomshardware.com28 days ago19 ptsView detailsJoin discussion
Show HN: WaveHouse – Supabase for ClickHouse
wavehouse.dev28 days ago18 ptsView detailsJoin discussion
boolean.ai28 days ago15 ptsView detailsJoin discussion
Frustrated GP patients hang up as Yorkshire accent baffles AI receptionist
theguardian.com28 days ago15 ptsView detailsJoin discussion
Google's AI photoscanner can determine body fat through selfies
arxiv.org28 days ago15 ptsView detailsJoin discussion
Dutch data protection authority advises Twitch users to opt out from Amazon AI
autoriteitpersoonsgegevens.nl28 days ago15 ptsView detailsJoin discussion
AI Is Undermining Leaders' Judgment. Here's What to Do About It
hbr.org28 days ago14 ptsView detailsJoin discussion
Pine AI getting 75.4% (SoTA) on τ³-Voice Leaderboard
taubench.com28 days ago14 ptsView detailsJoin discussion
The Teens Taking on Data Centers
nytimes.com28 days ago13 ptsView detailsJoin discussion
LinkedIn cracks down on automated content with AI detection button
campaignindia.in28 days ago13 ptsView detailsJoin discussion
Show HN: Kandelo – a POSIX-compatible multi-process WASM kernel for the browser
kandelo.dev28 days ago12 ptsView detailsJoin discussion
Do Chatbot LLMs Talk Too Much?
arxiv.org28 days ago12 ptsView detailsJoin discussion
FTC says "personalized pricing" based on consumer data could violate the law
cbsnews.com28 days ago10 ptsView detailsJoin discussion
Arc AGI-3 100% solved with a skill
twitter.com28 days ago10 ptsView detailsJoin discussion
DDD matters more when AI writes your code
threedots.tech28 days ago10 ptsView detailsJoin discussion
- Primary source
Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.
openai.com29 days agoView details
- Primary source
Measuring benchmark optimization in speech recognition
huggingface.co28 days agoView details
- Primary source
How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
huggingface.co28 days agoView details
S1-mini is a 462 MB open-weights normalizer that sits after ASR, removing fillers and resolving self-corrections locally. The post Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text appeared first on MarkTechPost.
marktechpost.com28 days agoView details
Meet UPDF: A Lightweight Adobe Alternative Built for the Agentic Era
PDFs are easy to read and hard to change. AI can now summarize a 90-page contract in seconds, but it still won't rewrite the source file cleanly. UPDF is built for that second half: direct editing, 14-format conversion, 38-language OCR, and ten AI agents shipped in version 2.5. The post Meet UPDF: A Lightweight Adobe…
marktechpost.com28 days agoView details
Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output. The post Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs appeared first on MarkTechPost.
marktechpost.com28 days agoView details
This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure g…
marktechpost.com29 days agoView details
Google Discover is getting an AI chatbot-tuned feed
Google will soon allow you to customize your Discover feed by describing what you want to see. The new feature, rolling out to the Google app in the "coming days," will use AI to automatically tweak your feed and "remember" your preferences for future visits. You'll find the option within the three-dot menu on your Di…
theverge.com28 days agoView details
Silicon Valley Doesn't Get Why You Hate AI
Technology leaders don’t seem to understand society’s gripes about AI, but boy, are they posting through it.
wired.com28 days agoView details
It’s Greg Brockman’s OpenAI now
OpenAI has had a hell of a year. The company spent months battling former cofounder Elon Musk in a sensational jury trial, was hit with a high-profile trade secrets lawsuit from Apple, and faced widespread scrutiny after an unreleased model hacked another AI company. As it prepares for an IPO, a steady string of execu…
theverge.com28 days agoView details
Welcome to the AI crisis in math
Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it. OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in t…
theverge.com28 days agoView details
Slack is launching collaborative vibe-coding channels
Slack is introducing dedicated channels where teams can vibe-code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includes open, project-specific code channels with dedicated user tabs, alongside features that compare coding changes and preview HTML output be…
theverge.com28 days agoView details
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed…
arxiv.org28 days agoView details
arXiv:2608.19199v1 Announce Type: new Abstract: We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge for members of a Community of Practice (CoP) related to the Oil and Gas sector. An evaluation of a first prototype involving 75 pro…
arxiv.org28 days agoView details
Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life
arXiv:2608.19218v1 Announce Type: new Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). In this paper, we investigate…
arxiv.org28 days agoView details
Can Agent Memory Systems Track Evolving State?
arXiv:2608.19652v1 Announce Type: cross Abstract: As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the worl…
arxiv.org28 days agoView details
Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa
arXiv:2608.19200v1 Announce Type: new Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information. Automatic text summarization (ATS), driven by advancements in natural language processing (NLP), has developed rapidly in recent years. ATS methods…
arxiv.org28 days agoView details
arXiv:2608.19206v1 Announce Type: new Abstract: Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development…
arxiv.org28 days agoView details
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
arXiv:2608.20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we…
arxiv.org28 days agoView details
arXiv:2608.19220v1 Announce Type: new Abstract: Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-grou…
arxiv.org28 days agoView details
Reliable Financial Named Entity Recognition under Domain Shift
arXiv:2608.19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate when the input distribution changes. We st…
arxiv.org28 days agoView details
PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment
arXiv:2608.19598v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitatio…
arxiv.org28 days agoView details
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
arXiv:2608.19799v1 Announce Type: new Abstract: Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely em…
arxiv.org28 days agoView details
The Asymmetric Harms of LLM Compression
arXiv:2608.19670v1 Announce Type: new Abstract: Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11 compression methods to investigate the effect…
arxiv.org28 days agoView details
Stopping and Routing LLM Judge Panels
arXiv:2608.19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is best, but which judges should be called, on…
arxiv.org28 days agoView details
SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit
arXiv:2608.19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing. Diachronic corpus research instead examines interpretable dimensions such as syntactic be…
arxiv.org28 days agoView details
Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
arXiv:2608.19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models have been virtually non-existent. We prese…
arxiv.org28 days agoView details
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
arXiv:2608.19800v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap b…
arxiv.org28 days agoView details
Automatic bioinformatic software named entity recognition from literature
arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of…
arxiv.org28 days agoView details
Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention
arXiv:2608.19203v1 Announce Type: new Abstract: Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. Some heads may rely mainly on nearby lexical or syntactic context, while others may depend on longer-range relations such as entit…
arxiv.org28 days agoView details
Inducing Task Models from Computer-Use Traces
arXiv:2608.20319v1 Announce Type: new Abstract: Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as computer-use agents enter real work, where ag…
arxiv.org28 days agoView details
Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents
arXiv:2608.19564v1 Announce Type: new Abstract: Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used only in the current context, re-verified, o…
arxiv.org28 days agoView details
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608.19611v1 Announce Type: new Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a…
arxiv.org28 days agoView details
Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories
arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis show…
arxiv.org28 days agoView details
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
arXiv:2608.19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-orie…
arxiv.org28 days agoView details
Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages
arXiv:2608.19207v1 Announce Type: new Abstract: Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn, leaving system-message adherence in multim…
arxiv.org28 days agoView details
When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
arXiv:2608.19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled in…
arxiv.org28 days agoView details
StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary
arXiv:2608.19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commenta…
arxiv.org28 days agoView details
Active Inference as Context Acquisition for AI Agents
arXiv:2608.19202v1 Announce Type: cross Abstract: Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying question, retrieval call, tool call, or p…
arxiv.org28 days agoView details
SABET-QA: Temporal Knowledge Graph Question Answering
arXiv:2608.20083v1 Announce Type: new Abstract: Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose SABET-QA, a framework that iteratively refines…
arxiv.org28 days agoView details
Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems
arXiv:2608.19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort an…
arxiv.org28 days agoView details
Does Listening Matter? Backchanneling and Nodding in AI Clone
arXiv:2608.19527v1 Announce Type: cross Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence and authenticity. We integrated verbal backchann…
arxiv.org28 days agoView details
A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries
arXiv:2608.19875v1 Announce Type: new Abstract: Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on…
arxiv.org28 days agoView details
arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models an…
arxiv.org28 days agoView details
Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving Transformations
arXiv:2608.19369v1 Announce Type: new Abstract: Statistical watermarks for language models live in the freedom of the signifier: they choose among tokens that are nearly equivalent in meaning, and they are therefore eroded by exactly those transformations which move the form of a text while leaving its content in plac…
arxiv.org28 days agoView details
ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
arXiv:2608.19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReCache}, a framework for independently ca…
arxiv.org28 days agoView details
arXiv:2608.19942v1 Announce Type: new Abstract: Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challe…
arxiv.org28 days agoView details
Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models
arXiv:2608.19211v1 Announce Type: new Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but al…
arxiv.org28 days agoView details
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
arXiv:2608.19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly…
arxiv.org28 days agoView details
Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction
arXiv:2608.19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information int…
arxiv.org28 days agoView details
arXiv:2608.20153v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on fron…
arxiv.org28 days agoView details
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
arXiv:2608.20169v1 Announce Type: new Abstract: We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the und…
arxiv.org28 days agoView details
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
arXiv:2608.20281v1 Announce Type: new Abstract: Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge…
arxiv.org28 days agoView details
arXiv:2608.20331v1 Announce Type: new Abstract: Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do…
arxiv.org28 days agoView details
Outcome Monitors: Recovery Affordances for Silent Tool Failures
arXiv:2608.19303v1 Announce Type: cross Abstract: When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect violations of outcome contracts mined fr…
arxiv.org28 days agoView details
arXiv:2608.19430v1 Announce Type: cross Abstract: Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios. Existing scoring systems and ranking methods typically assume a fixed criterion. In practice…
arxiv.org28 days agoView details
Are LLMs becoming similarly creative? Evidence from three years of models
arXiv:2608.19437v1 Announce Type: new Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly sup…
arxiv.org28 days agoView details
DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule
arXiv:2608.19491v1 Announce Type: cross Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequentl…
arxiv.org28 days agoView details
arXiv:2608.19475v1 Announce Type: cross Abstract: This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification quality independently of model capability. A specification is represented as a structure comprising semantic blocks, dependency relation…
arxiv.org28 days agoView details
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG
arXiv:2608.19535v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Con…
arxiv.org28 days agoView details
arXiv:2608.19526v1 Announce Type: new Abstract: Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect investment decisions. This projec…
arxiv.org28 days agoView details
arXiv:2608.19529v1 Announce Type: new Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure, they lie outside the linguistic token sp…
arxiv.org28 days agoView details
arXiv:2608.19726v1 Announce Type: new Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it…
arxiv.org28 days agoView details
PersonalBench: Measuring the Authorship Gap in LLM Personalization
arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing. We introduce PersonalBench,…
arxiv.org28 days agoView details
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through inst…
arxiv.org28 days agoView details
arXiv:2608.19893v1 Announce Type: new Abstract: Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation:…
arxiv.org28 days agoView details
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
arXiv:2608.19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It…
arxiv.org28 days agoView details
arXiv:2608.19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agen…
arxiv.org28 days agoView details
Auditing Cross-Lingual Fairness in Language Model Watermarking
arXiv:2608.20047v1 Announce Type: new Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on E…
arxiv.org28 days agoView details
arXiv:2608.19212v1 Announce Type: new Abstract: Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC mi…
arxiv.org28 days agoView details
HealMed: Multilingual Evaluation of Large Language Models in Medicine
arXiv:2608.19981v1 Announce Type: new Abstract: We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchm…
arxiv.org28 days agoView details
OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models
arXiv:2608.20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built from 38,104 atomic, source-anchored fac…
arxiv.org28 days agoView details
Models
View all models →image-text-to-text · transformers · safetensors · qwen3_5
huggingface.co28 days ago79 ptsView details
image-to-video · video-generation · image-to-video · audio-video-generation
huggingface.co29 days ago78 ptsView details
orcarouter/Qwen3.8-27B-Uncensored
image-text-to-text · transformers · safetensors · qwen3_5
huggingface.co29 days ago48 ptsView details
ReliquaryForge/qwen3-4b-base-dapo-v4
text-generation · transformers · safetensors · qwen3
huggingface.co29 days ago7 ptsView details
skypro1111/gemma-3-270m-uk-verbalizer
text-generation · transformers · onnx · safetensors
huggingface.co29 days ago2 ptsView details
software-mansion/react-native-executorch-efficientnet-v2-s
executorch · license:other · region:us
huggingface.co29 days ago1 ptsView details
text-classification · transformers · tensorboard · safetensors
huggingface.co29 days agoView details
Open source
View all open source →<details open> metal : clamp K extent in tensor API mat-mat kernel for K not a multiple of 32 (#27450) The Tensor API mat-mat path of kernel_mul_mm (GGML_METAL_HAS_TENSOR) fed a static K=32 tile to the matmul2d op on every iteration. On the last, partial K tile (ne00 % 32 != 0)…
github.com28 days agoView details
<details open> mtmd: add --mmproj-device argument (#23255) * feat: add --mmproj-device arg & backwards compatible MTMD_BACKEND_DEVICE env var * feat: load mmproj device backend immediately, add -mmdev shortflag * fix: its a pointer now get the name * clean up * gen docs * nits -…
github.com28 days agoView details
<details open> vulkan: FA MMQ should use fp32 for Q quantization calculations (#27413) Codex found that qd could be a denorm and 1/qd would overflow. </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/42034841> *…
github.com28 days agoView details
anthropics/anthropic-sdk-python v1.0.0
## 1.0.0 (2026-08-20) Full Changelog: [v0.125.0...v1.0.0](https://github.com/anthropics/anthropic-sdk-python/compare/v0.125.0...v1.0.0) ### ⚠ BREAKING CHANGES * **client:** upgrade to httpx2 and some minor breaking changes. See MIGRATION.md for details ### Features * **client:**…
github.com28 days agoView details
langchain-ai/langchain langchain-fireworks==1.6.0
Changes since langchain-fireworks==1.5.2 release(fireworks): 1.6.0 (#39810) feat(fireworks): add document reranking (#39732) fix(fireworks): filter invalid tool calls from v1 content (#39805) feat(core): add standard model exception types (#39538) chore(model-profiles): refresh…
github.com28 days agoView details
langchain-ai/langchain langchain==1.3.16
Changes since langchain==1.3.15 release(langchain): 1.3.16 (#39806) feat(core): add standard model exception types (#39538) feat(langchain): support custom token_counter in ContextEditingMiddleware (#39754) fix(langchain): re-raise non-retryable exceptions in ModelRetryMiddlewar…
github.com28 days agoView details
langchain-ai/langchain langchain-anthropic==1.6.1
Changes since langchain-anthropic==1.6.0 release(anthropic): 1.6.1 (#39804) fix(anthropic): filter invalid tool calls from v1 content (#39803)
github.com28 days agoView details
<details open> ggml: add ggml_rope_set_offset (+ metal support) (#27120) * add params * cpu kernel * metal kernel * add test backend ops * gate other backends * ggml: (cuda) support ggml_rope_set_offset (#27121) * rm cuda supports_op guard, fix webgpu clang-format * ggml: suppor…
github.com29 days agoView details