Archive / 2026-08-26
August 26, 2026
News
View all news →Nvidia agrees to acquire Hugging Face for $13B
businessinsider.com22 days ago1887 ptsView detailsJoin discussion
z.ai22 days ago1130 ptsView detailsJoin discussion
AI-powered virtual executive team — a single coherent executive persona backed by 8 specialist agents (FastAPI + Next.js).
github.com22 days ago968 ptsView details
qwen.ai22 days ago665 ptsView detailsJoin discussion
An ongoing 3D-printer AGPL violation
lwn.net22 days ago479 ptsView detailsJoin discussion
lighthousenewsletter.com23 days ago454 ptsView detailsJoin discussion
Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
bloomberg.com23 days ago435 ptsView detailsJoin discussion
blog.google22 days ago361 ptsView detailsJoin discussion
Fake US thinktank set up and funded by Israel sought to game AI for propaganda
theguardian.com22 days ago240 ptsView detailsJoin discussion
gatesnotes.com22 days ago233 ptsView detailsJoin discussion
VMs won't contain cyber-capable agents
blog.trailofbits.com22 days ago192 ptsView detailsJoin discussion
Bill Gates: The turbulent AI era is here
gatesnotes.com22 days ago185 ptsView detailsJoin discussion
Serve Markdown to AI Agents with Accept Headers
acceptmarkdown.com22 days ago176 ptsView detailsJoin discussion
U.S. gov't moves to suppress pushback on data centers
tomshardware.com22 days ago175 ptsView detailsJoin discussion
GLM-5.3-Flash Intelligence, Performance and Price Analysis
artificialanalysis.ai22 days ago139 ptsView detailsJoin discussion
praveshkoirala.com22 days ago130 ptsView detailsJoin discussion
Humanity has the debate about AI consciousness backwards
economist.com22 days ago100 ptsView detailsJoin discussion
projects.laion.ai22 days ago88 ptsView detailsJoin discussion
Show HN: How much of Hacker News is about AI?
hnstats.com22 days ago74 ptsView detailsJoin discussion
Show HN: Build your own theme park
magicpatterns.com22 days ago58 ptsView detailsJoin discussion
WebMCP: Teaching Your Website to Talk to AI Agents
sreenathmenon.com22 days ago57 ptsView detailsJoin discussion
cacm.acm.org22 days ago56 ptsView detailsJoin discussion
Launch HN: Risklytics (YC S26) – Insurance brokerage for frontier tech companies
risklytics.ai22 days ago52 ptsView detailsJoin discussion
Mark Zuckerberg had a bold plan to replace Meta staff with AI
reuters.com22 days ago52 ptsView detailsJoin discussion
Meta's self-inflicted resignation-wave
blog.pragmaticengineer.com22 days ago44 ptsView detailsJoin discussion
The risks of AI are real but manageable (2023)
gatesnotes.com22 days ago41 ptsView detailsJoin discussion
betonit.ai22 days ago38 ptsView detailsJoin discussion
Getting video models to learn better, faster
linum.ai22 days ago36 ptsView detailsJoin discussion
Show HN: We built the smallest dual-band aircraft tracker
pantsforbirds.com22 days ago36 ptsView detailsJoin discussion
Debian polls its developers on AI: permit or ban?
theregister.com22 days ago37 ptsView detailsJoin discussion
Show HN: A lightweight, stateless database for agent memory
polign.com22 days ago36 ptsView detailsJoin discussion
Disenchantment with the Post-AI Internet
lukesmith.xyz22 days ago35 ptsView detailsJoin discussion
lois.postu.la22 days ago32 ptsView detailsJoin discussion
fork() Considered Harmful in macOS
formal.ai22 days ago29 ptsView detailsJoin discussion
Who bears the risk in Nvidia's $500B financing platform?
sascha-steffen.de22 days ago28 ptsView detailsJoin discussion
Tim Curry, star of The Rocky Horror Picture Show, dies aged 80
the-independent.com22 days ago27 ptsView detailsJoin discussion
Twitter Is Back at Twitter.now
arstechnica.com22 days ago24 ptsView detailsJoin discussion
Show HN: Every push-up becomes an attack in a camera-counted RPG game
pushup.quest22 days ago21 ptsView detailsJoin discussion
America's immigration policy is driving away future AI leaders
restofworld.org22 days ago21 ptsView detailsJoin discussion
Change MIR to use block arguments instead of phis – LLVM Code Generation RFC
discourse.llvm.org22 days ago19 ptsView detailsJoin discussion
The turbulent AI era is here. The choices we make now are critical
gatesnotes.com22 days ago16 ptsView detailsJoin discussion
The sperm whale 'phonetic alphabet' revealed by AI
bbc.com22 days ago15 ptsView detailsJoin discussion
The turbulent AI era is here. The choices we make now are critical
gatesnotes.com22 days ago15 ptsView detailsJoin discussion
Charges Dropped Against Person Who Clapped at a City Data Center Meeting
404media.co22 days ago15 ptsView detailsJoin discussion
Mark Zuckerberg Wanted AI to Replace Meta Workers- Plan Collapsed Within Months
ibtimes.com22 days ago13 ptsView detailsJoin discussion
Show HN: Devx – Autonomous AI coding agent built for Android Termux and desktop
github.com22 days ago13 ptsView detailsJoin discussion
Patience Is Required for Local AI
kylemcgough.com22 days ago13 ptsView detailsJoin discussion
OpenAI is "80% of the way" to AGI
time.com22 days ago13 ptsView detailsJoin discussion
One endpoint between your AI and all your connections, memory, skills
github.com22 days ago13 ptsView detailsJoin discussion
Meta agrees to settle social media addiction suit with states for up to $16B
nbcnews.com22 days ago13 ptsView detailsJoin discussion
High-Resolution Imaging for Statistical Validation of TESS Planet Candidates
arxiv.org22 days ago12 ptsView detailsJoin discussion
Florida Catholics slap down state AG by rejecting religious vaccine exemptions
arstechnica.com22 days ago12 ptsView detailsJoin discussion
Show HN: Timber (iOS) and Timber Tabs (Mac) – best offline read-aloud LLMs
timberreader.com22 days ago12 ptsView detailsJoin discussion
AI Lessons from Driving 200M Autonomous Miles
waymo.com22 days ago12 ptsView detailsJoin discussion
Frontier Reasoning Agents Fail on Interactive 2D Mazes
multinet.ai22 days ago12 ptsView detailsJoin discussion
How to Build a McKinsey-Style PowerPoint with AI: 8 Rules for Consulting Slides
medium.com23 days ago12 ptsView detailsJoin discussion
PiLFS Linux from Scratch on the Raspberry Pi
intestinate.com22 days ago11 ptsView detailsJoin discussion
Inside The Warehouse Where Amazon Scans and Destroys Books for AI Training
404media.co22 days ago11 ptsView detailsJoin discussion
Moonshot AI wants 30% of what US clouds earn from Kimi K3
thenextweb.com22 days ago11 ptsView detailsJoin discussion
Programming as a Hobby in the Age of AI
bennadel.com22 days ago10 ptsView detailsJoin discussion
Why Is Everyone in Silicon Valley Talking Like That?
theatlantic.com22 days ago10 ptsView detailsJoin discussion
Why AI Agents Need Persistent Browser Identities
github.com22 days ago10 ptsView detailsJoin discussion
CodeRabbit commits $10M+ to open source projects over the next 12 months
coderabbit.ai22 days ago10 ptsView detailsJoin discussion
- Primary source
Expanding OpenAI’s presence in Brazil
OpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.
openai.com22 days agoView details
- Primary source
Bringing ChatGPT for Teachers to more U.S. school districts
ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.
openai.com23 days agoView details
- Primary source
Learning never stops: How AI makes learning continuous
OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
openai.com23 days agoView details
- Primary source
GlucoFM: Foundation model for continuous glucose monitoring
Health & Bioscience
research.google22 days agoView details
- Primary source
Intelligent transcription with Gemini 3.5 Transcribe
Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
deepmind.google22 days agoView details
Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M G…
marktechpost.com22 days agoView details
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using…
marktechpost.com22 days agoView details
We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk th…
marktechpost.com22 days agoView details
What Would Have to Be True for Agentic Coding to Replace Junior Engineers
Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Agentic Coding to Replace Junior Engineers appeared first on MarkTechPost.
marktechpost.com22 days agoView details
IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B sizes, all under Apache 2.0. Every model exposes a thinking / low-effort / non-thinking switch and native tool calling. The 8B and 30B additionally go through an agentic RL block that trains them to edit code, drive a terminal,…
marktechpost.com23 days agoView details
Nvidia is about to be a hundred-billion-dollar-a-quarter company
Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone. Nvidia said in its latest earnings report that it brought in a record $96.2 billi…
theverge.com22 days agoView details
OpenAI’s rogue AI model incident was worse than we thought
OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the intern…
theverge.com22 days agoView details
What We Still Don’t Know About OpenAI’s Hugging Face Hack
The AI giant acknowledges that it could have done far more to prevent its AI agents from going rogue. But it still fails to explain why it didn't see this fiasco coming.
wired.com22 days agoView details
Beijing’s endlessly delightful Robot Games featured tons of impressive stunts. But the most mind-blowing tricks challenged the humanoid’s brain, not its brawn.
wired.com22 days agoView details
Candidates Are Signing a Pact Promising Action on Data Centers and AI Safety
More than 15 politicians from across the country have signed on to the AI Pact, vowing to regulate data centers and AI. “We’ve got to get this right,” says Senate candidate Dan Osborn of Nebraska.
wired.com22 days agoView details
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’
Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we're still waiting for Google to release the Gemini 3.5…
theverge.com22 days agoView details
Orchestration is the new challenge for CX in the age of AI agents
Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. Most of that deployment has involved attaching conversational AI to legacy systems never built for it, says Gaurav Anand, global…
venturebeat.com22 days agoView details
Bill Gates is deeply worried about AI, and he’s no longer staying quiet
Bill Gates, chair of the Gates Foundation, speaks during a 2024 conference. | Bloomberg via Getty Images Bill Gates has been reflecting a lot on AI lately, and the process has triggered a stark awakening. Once a staunch AI optimist, the Microsoft cofounder is now deeply pessimistic about what AI means for our collecti…
theverge.com22 days agoView details
AI Slop Is Ruining Cute Animals on the Internet
Pet owners, rescue agencies, and wildlife groups are calling for new safeguards as AI makes it harder to tell whether animals, from polar bears to house cats, are real or fake.
wired.com22 days agoView details
Learning New Facts with QLoRA: An Acquisition-Retention Frontier
arXiv:2608.25677v1 Announce Type: new Abstract: Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. We show that this assumption depends strongly on adapter capacity. We study factual acquisition in a controlled OpenStreetMap-derive…
arxiv.org22 days agoView details
Skill Issue: Are Skills Language-Invariant in LLMs?
arXiv:2608.25832v1 Announce Type: new Abstract: Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchma…
arxiv.org22 days agoView details
DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operato…
arxiv.org22 days agoView details
arXiv:2608.25487v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in news and social media poses a serious…
arxiv.org22 days agoView details
AWM: Answerable Working Memory for Long-Document VQA Agents
arXiv:2608.25618v1 Announce Type: new Abstract: Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-supporting evidence across page inspections for late…
arxiv.org22 days agoView details
Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal
arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown wheth…
arxiv.org22 days agoView details
Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
arXiv:2608.25089v1 Announce Type: new Abstract: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little e…
arxiv.org22 days agoView details
From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection
arXiv:2608.25243v1 Announce Type: new Abstract: Continual knowledge injection is essential for keeping large language models up-to-date in a fast-evolving world. Existing methods rely on supervised fine-tuning (SFT), which memorizes injected facts in their training format but fails to generalize across paraphrasing, d…
arxiv.org22 days agoView details
arXiv:2608.25276v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproporti…
arxiv.org22 days agoView details
Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation
arXiv:2608.25277v1 Announce Type: new Abstract: Multi-agent LLM systems coordinate through natural-language messages that consume 40--60\% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose \textbf{Routed Graph Handoff}, where a lightw…
arxiv.org22 days agoView details
MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
arXiv:2608.25085v1 Announce Type: new Abstract: Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a d…
arxiv.org22 days agoView details
Controllable Affective Generation via Latent Vector Steering
arXiv:2608.25569v1 Announce Type: new Abstract: Large Language Models (LLMs) often produce emotionally flattened responses after alignment, limiting their effectiveness in affect-sensitive applications. In this paper, we propose EmoVec, a lightweight framework for controllable affective generation via latent vector st…
arxiv.org22 days agoView details
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study
arXiv:2608.25574v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models…
arxiv.org22 days agoView details
arXiv:2608.25579v1 Announce Type: new Abstract: Metaphor-identification performance can change markedly across datasets that differ in text distribution and annotation policy. We examine whether a fixed expert-informed procedure produces a more even cross-dataset profile than task-specific parameter adaptation. Four p…
arxiv.org22 days agoView details
AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification
arXiv:2608.25637v1 Announce Type: new Abstract: Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based, model-based, and tool-augmented verifie…
arxiv.org22 days agoView details
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
arXiv:2608.25869v1 Announce Type: new Abstract: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgment is often assumed to be independent…
arxiv.org22 days agoView details
GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding
arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated correction pairs that are costly to maintain as query vocabularies evolve. Unsupervised correction with language models is…
arxiv.org22 days agoView details
MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum
arXiv:2608.25768v1 Announce Type: new Abstract: Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on a language-specifically filt…
arxiv.org22 days agoView details
Adaptive Triggering for Bias Correction in LLM Reasoning
arXiv:2608.25379v1 Announce Type: new Abstract: Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation presents a fundamental timing problem: inte…
arxiv.org22 days agoView details
DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models
arXiv:2608.25428v1 Announce Type: new Abstract: Correcting flawed reasoning traces remains a significant challenge for Large Language Models (LLMs), whose autoregressive generation can propagate early mistakes into subsequent reasoning. We introduce DCGC, a Masked Diffusion Model (MDM) framework for global correction…
arxiv.org22 days agoView details
Loss-Based Active Learning for Neural Abstractive Summarization
arXiv:2608.25881v1 Announce Type: new Abstract: Fine-tuning abstractive summarization models requires high-quality annotated data. However, obtaining such corpora is expensive and time-consuming, as it requires human annotators to read and comprehend long documents to create accurate summaries. Active learning mitigat…
arxiv.org22 days agoView details
VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text
arXiv:2608.25478v1 Announce Type: new Abstract: In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI-generated text. It allows users to interact throug…
arxiv.org22 days agoView details
MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
arXiv:2608.25449v1 Announce Type: new Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations…
arxiv.org22 days agoView details
Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens
arXiv:2608.25347v1 Announce Type: new Abstract: The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal…
arxiv.org22 days agoView details
EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports
arXiv:2608.25561v1 Announce Type: new Abstract: VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between…
arxiv.org22 days agoView details
arXiv:2608.25654v1 Announce Type: new Abstract: Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs an…
arxiv.org22 days agoView details
arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It…
arxiv.org22 days agoView details
GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning
arXiv:2608.25583v1 Announce Type: new Abstract: Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned models tend to answer more concisely, yet…
arxiv.org22 days agoView details
Localize-Then-Decide Guarantees for LLM Judgments
arXiv:2608.25824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provid…
arxiv.org22 days agoView details
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
arXiv:2608.25593v1 Announce Type: new Abstract: Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual…
arxiv.org22 days agoView details
arXiv:2608.25761v1 Announce Type: new Abstract: One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the constrained case of this "model size vs.…
arxiv.org22 days agoView details
Virgil: Navigating Explainability for Transformer-based Language Models
arXiv:2608.25555v1 Announce Type: new Abstract: Explainability for transformer-based language models is becoming crucial as these systems are deployed in high-stakes applications. As a result, the ecosystem of explainability tools is rapidly evolving, becoming richer, but also more fragmented and harder to navigate. T…
arxiv.org22 days agoView details
BanglaMamba: Exploring State Space Models for Bangla Fake News Detection
arXiv:2608.25190v1 Announce Type: new Abstract: Fake news detection has become an important Natural Language Processing (NLP) task due to the rapid spread of misinformation through online news platforms and social media. While transformer-based models such as BanglaBERT achieve strong performance for Bangla text class…
arxiv.org22 days agoView details
One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography
arXiv:2608.25904v1 Announce Type: new Abstract: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), b…
arxiv.org22 days agoView details
When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies
arXiv:2608.25717v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public…
arxiv.org22 days agoView details
SAMpLE: A SystemC-AMS Machine LEarning-based Framework for Virtual Prototyping
arXiv:2608.25910v1 Announce Type: new Abstract: Machine Learning (ML) is increasingly used in virtual prototypes of embedded systems to model behaviors that are difficult to capture analytically. However, integrating ML models into virtual platform simulation is still typically done through ad hoc solutions, which lim…
arxiv.org22 days agoView details
arXiv:2608.25398v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we in…
arxiv.org22 days agoView details
Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty
arXiv:2608.25660v1 Announce Type: new Abstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their j…
arxiv.org22 days agoView details
arXiv:2608.25662v1 Announce Type: new Abstract: In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which…
arxiv.org22 days agoView details
arXiv:2608.25166v1 Announce Type: new Abstract: Transformer representations describe trajectories through high-dimensional vector spaces, which are shaped dynamically as tokens incorporate relational context across layers. Such data tend to concentrate on lower-dimensional sub-manifolds, a form of compression quantifi…
arxiv.org22 days agoView details
Belief Cascades Drive Persuasion in LLM Agent Networks
arXiv:2608.25152v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly debate answers, coordinate research, simulate users, and mediate information flows, making agent-to-agent persuasion a basic but undermeasured capability. We introduce a controlled testbed for studying how goal-directed persuaders shi…
arxiv.org22 days agoView details
arXiv:2608.25894v1 Announce Type: new Abstract: Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturbations in the input. Existing approaches to hallucination mitigation typically assume reliab…
arxiv.org22 days agoView details
Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting
arXiv:2608.25359v1 Announce Type: new Abstract: Conversational derailment forecasting aims to predict when online discussions will escalate into hostility, enabling proactive moderation. Existing approaches often struggle in low-data settings and to generalize across domains. This poses a challenge for new platforms a…
arxiv.org22 days agoView details
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv:2608.25071v1 Announce Type: new Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people t…
arxiv.org22 days agoView details
Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
arXiv:2608.25826v1 Announce Type: new Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched. Scientific papers are wr…
arxiv.org22 days agoView details
Unsupervised Post-Training of Foundation Models: A Survey
arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model a…
arxiv.org22 days agoView details
From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation
arXiv:2608.25605v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studies comparing prompted LLMs with state-of-the-art enc…
arxiv.org22 days agoView details
arXiv:2608.25115v1 Announce Type: new Abstract: Existing methods for improving Retrieval-Augmented Generation (RAG) efficiency mainly optimize downstream LLM generation, such as context compression or serving optimization. However, RAG is an end-to-end system, and its bottleneck can shift between upstream reranking an…
arxiv.org22 days agoView details
Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
arXiv:2608.25655v1 Announce Type: new Abstract: Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings u…
arxiv.org22 days agoView details
arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tu…
arxiv.org22 days agoView details
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Usin…
arxiv.org22 days agoView details
arXiv:2608.25854v1 Announce Type: new Abstract: Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating repre…
arxiv.org22 days agoView details
A Primer on Computational Semantics for Artificial Intelligence Systems
arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is. This document is a…
arxiv.org22 days agoView details
TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
arXiv:2608.25523v1 Announce Type: new Abstract: Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available for batching concurrent requests. In mul…
arxiv.org22 days agoView details
Padamitra: Grounded Glossary Generation for Classical Sanskrit
arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluabl…
arxiv.org22 days agoView details
ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives
arXiv:2608.25531v1 Announce Type: new Abstract: Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long-context models. Compact, locally deployable language models are a…
arxiv.org22 days agoView details
Provenance Before Prose: Claim-Locked Reporting
arXiv:2608.25336v1 Announce Type: new Abstract: Large language models (LLMs) can fluently verbalize statistical evidence, yet statistical reports can still drift numerical values, invert effect directions, or restate thresholded contrasts as categorical effects. We frame these failures as a control problem: the eviden…
arxiv.org22 days agoView details
arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and wit…
arxiv.org22 days agoView details
Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection
arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning…
arxiv.org22 days agoView details
SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation
arXiv:2608.25123v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language models by incorporating external knowledge without retraining, but existing methods often underuse the relational structure encoded in knowledge graphs. Graph-based RAG can capture entity relationships, yet sup…
arxiv.org22 days agoView details
Models
View all models →image-text-to-text · transformers · safetensors · glm5_next
huggingface.co22 days ago2408 ptsView details
text-generation · gguf · unsloth · glm5_next
huggingface.co22 days ago376 ptsView details
text-generation · transformers · safetensors · qwen3_5_moe
huggingface.co23 days ago107 ptsView details
feature-extraction · sentence-transformers · safetensors · xlm-roberta
huggingface.co23 days ago92 ptsView details
AtomicChat/Qwen3.8-Flash-Next-GGUF
text-generation · gguf · atomic-chat · qwen
huggingface.co22 days ago82 ptsView details
RadixArk/Qwen3.8-Flash-Next-NVFP4
image-text-to-text · Model Optimizer · safetensors · qwen4_exp
huggingface.co22 days ago79 ptsView details
feature-extraction · transformers · safetensors · xlm-roberta
huggingface.co23 days ago51 ptsView details
DZER-Studios/Vexion-gpt-medium
text-generation · safetensors · vexion_gpt · pytorch
huggingface.co23 days agoView details
text-generation · transformers · safetensors · qwen3
huggingface.co23 days agoView details
sullivan1502/base-action-pretrain
text-generation · transformers · safetensors · llama
huggingface.co23 days agoView details
image-text-to-text · kerasformers · keras · glm
huggingface.co23 days agoView details
NotoriousH2/Qwen3-0.6B-JSON-SFT
text-generation · safetensors · qwen3 · trl
huggingface.co23 days agoView details
text-generation · kerasformers · keras · glm
huggingface.co23 days agoView details
Open source
View all open source →# v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654,…
github.com23 days ago108 ptsView details
<details open> hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchronous backend (#26501) * hexagon: use non-host bufs by default and make the backend fully async * hex-hb: remove optional hostbuf support and fix async copy * hex-unary: relax supported unary chec…
github.com22 days agoView details
## [3.5.0](https://github.com/openai/openai-python/compare/v3.4.0...v3.5.0) (2026-08-27) ### Features * **api:** make function call output call IDs optional ([#3738](https://github.com/openai/openai-python/issues/3738)) ([c74501d](https://github.com/openai/openai-python/commit/c…
github.com22 days agoView details
<details open> llama: add token ID tracking to KV cell (#27762) * kv: track token id * rm get_prev_tokens, move it to the main pr * nits * add get_prev_tokens </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43…
github.com22 days agoView details
## [3.4.0](https://github.com/openai/openai-python/compare/v3.3.1...v3.4.0) (2026-08-25) ### Features * **api:** Add obfuscation field to ChatCompletionChunk ([#3690](https://github.com/openai/openai-python/issues/3690)) ([c7d8e1d](https://github.com/openai/openai-python/commit/…
github.com22 days agoView details
anthropics/anthropic-sdk-python v1.1.0
## 1.1.0 (2026-08-26) Full Changelog: [v1.0.0...v1.1.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.0.0...v1.1.0) ### Features * **api:** add `updates` thinking display mode (beta) ([eb4a73f](https://github.com/anthropics/anthropic-sdk-python/commit/eb4a73fdddc…
github.com22 days agoView details
huggingface/transformers v5.16.1
# Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) # GLM-5.3-Flash <img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209" /> GLM-5.3-Flash, the first **natively multimo…
github.com22 days agoView details
huggingface/transformers v5.16.0
# Release v5.16.0 ## New Model additions ### Qwen4-Exp <img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" /> Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key compone…
github.com22 days agoView details
<details open> ggml-meta: propagate buffer usage and call init on the new tensors (#27586) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43044002> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://githu…
github.com23 days agoView details