Skip to content

Archive / 2026-08-26

August 26, 2026

  1. GLM-5.3-Flash

    z.ai22 days ago1130 ptsView detailsJoin discussion

  2. SenteLabsAI/OpenExecutive

    AI-powered virtual executive team — a single coherent executive persona backed by 8 specialist agents (FastAPI + Next.js).

    github.com22 days ago968 ptsView details

  3. Qwen3.8-Flash-Next

    qwen.ai22 days ago665 ptsView detailsJoin discussion

  4. RAG Is Simpler Than You Think

    lighthousenewsletter.com23 days ago454 ptsView detailsJoin discussion

  5. Gemini-3.5-Transcribe

    blog.google22 days ago361 ptsView detailsJoin discussion

  6. The turbulent AI era is here

    gatesnotes.com22 days ago233 ptsView detailsJoin discussion

  7. VMs won't contain cyber-capable agents

    blog.trailofbits.com22 days ago192 ptsView detailsJoin discussion

  8. PageRank explained

    praveshkoirala.com22 days ago130 ptsView detailsJoin discussion

  9. Laion Big Video Dataset

    projects.laion.ai22 days ago88 ptsView detailsJoin discussion

  10. Show HN: Build your own theme park

    magicpatterns.com22 days ago58 ptsView detailsJoin discussion

  11. AI Is a Harsh Mistress

    cacm.acm.org22 days ago56 ptsView detailsJoin discussion

  12. Meta's self-inflicted resignation-wave

    blog.pragmaticengineer.com22 days ago44 ptsView detailsJoin discussion

  13. CDs vs. NIMBY

    betonit.ai22 days ago38 ptsView detailsJoin discussion

  14. Infrastructure as Raclette

    lois.postu.la22 days ago32 ptsView detailsJoin discussion

  15. Twitter Is Back at Twitter.now

    arstechnica.com22 days ago24 ptsView detailsJoin discussion

  16. Patience Is Required for Local AI

    kylemcgough.com22 days ago13 ptsView detailsJoin discussion

  17. Primary source

    Expanding OpenAI’s presence in Brazil

    OpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.

    openai.com22 days agoView details

  18. Primary source

    Bringing ChatGPT for Teachers to more U.S. school districts

    ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.

    openai.com23 days agoView details

  19. Primary source

    Learning never stops: How AI makes learning continuous

    OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.

    openai.com23 days agoView details

  20. Primary source

    GlucoFM: Foundation model for continuous glucose monitoring

    Health & Bioscience

    research.google22 days agoView details

  21. Primary source

    Intelligent transcription with Gemini 3.5 Transcribe

    Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

    deepmind.google22 days agoView details

  22. Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring

    Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M G…

    marktechpost.com22 days agoView details

  23. Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

    Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using…

    marktechpost.com22 days agoView details

  24. Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

    We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk th…

    marktechpost.com22 days agoView details

  25. What Would Have to Be True for Agentic Coding to Replace Junior Engineers

    Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Agentic Coding to Replace Junior Engineers appeared first on MarkTechPost.

    marktechpost.com22 days agoView details

  26. IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models

    IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B sizes, all under Apache 2.0. Every model exposes a thinking / low-effort / non-thinking switch and native tool calling. The 8B and 30B additionally go through an agentic RL block that trains them to edit code, drive a terminal,…

    marktechpost.com23 days agoView details

  27. Nvidia is about to be a hundred-billion-dollar-a-quarter company

    Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone. Nvidia said in its latest earnings report that it brought in a record $96.2 billi…

    theverge.com22 days agoView details

  28. OpenAI’s rogue AI model incident was worse than we thought

    OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the intern…

    theverge.com22 days agoView details

  29. What We Still Don’t Know About OpenAI’s Hugging Face Hack

    The AI giant acknowledges that it could have done far more to prevent its AI agents from going rogue. But it still fails to explain why it didn't see this fiasco coming.

    wired.com22 days agoView details

  30. The Humanoids at China’s Robot Games Were Faster Than Usain Bolt—but I’m More Impressed by Their Tweezer Mastery

    Beijing’s endlessly delightful Robot Games featured tons of impressive stunts. But the most mind-blowing tricks challenged the humanoid’s brain, not its brawn.

    wired.com22 days agoView details

  31. Candidates Are Signing a Pact Promising Action on Data Centers and AI Safety

    More than 15 politicians from across the country have signed on to the AI Pact, vowing to regulate data centers and AI. “We’ve got to get this right,” says Senate candidate Dan Osborn of Nebraska.

    wired.com22 days agoView details

  32. Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

    Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we're still waiting for Google to release the Gemini 3.5…

    theverge.com22 days agoView details

  33. Orchestration is the new challenge for CX in the age of AI agents

    Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. Most of that deployment has involved attaching conversational AI to legacy systems never built for it, says Gaurav Anand, global…

    venturebeat.com22 days agoView details

  34. Bill Gates is deeply worried about AI, and he’s no longer staying quiet

    Bill Gates, chair of the Gates Foundation, speaks during a 2024 conference. | Bloomberg via Getty Images Bill Gates has been reflecting a lot on AI lately, and the process has triggered a stark awakening. Once a staunch AI optimist, the Microsoft cofounder is now deeply pessimistic about what AI means for our collecti…

    theverge.com22 days agoView details

  35. AI Slop Is Ruining Cute Animals on the Internet

    Pet owners, rescue agencies, and wildlife groups are calling for new safeguards as AI makes it harder to tell whether animals, from polar bears to house cats, are real or fake.

    wired.com22 days agoView details

  36. Learning New Facts with QLoRA: An Acquisition-Retention Frontier

    arXiv:2608.25677v1 Announce Type: new Abstract: Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. We show that this assumption depends strongly on adapter capacity. We study factual acquisition in a controlled OpenStreetMap-derive…

    arxiv.org22 days agoView details

  37. Skill Issue: Are Skills Language-Invariant in LLMs?

    arXiv:2608.25832v1 Announce Type: new Abstract: Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchma…

    arxiv.org22 days agoView details

  38. DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

    arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operato…

    arxiv.org22 days agoView details

  39. ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

    arXiv:2608.25487v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in news and social media poses a serious…

    arxiv.org22 days agoView details

  40. AWM: Answerable Working Memory for Long-Document VQA Agents

    arXiv:2608.25618v1 Announce Type: new Abstract: Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-supporting evidence across page inspections for late…

    arxiv.org22 days agoView details

  41. Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

    arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown wheth…

    arxiv.org22 days agoView details

  42. Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

    arXiv:2608.25089v1 Announce Type: new Abstract: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little e…

    arxiv.org22 days agoView details

  43. From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

    arXiv:2608.25243v1 Announce Type: new Abstract: Continual knowledge injection is essential for keeping large language models up-to-date in a fast-evolving world. Existing methods rely on supervised fine-tuning (SFT), which memorizes injected facts in their training format but fails to generalize across paraphrasing, d…

    arxiv.org22 days agoView details

  44. Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

    arXiv:2608.25276v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproporti…

    arxiv.org22 days agoView details

  45. Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation

    arXiv:2608.25277v1 Announce Type: new Abstract: Multi-agent LLM systems coordinate through natural-language messages that consume 40--60\% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose \textbf{Routed Graph Handoff}, where a lightw…

    arxiv.org22 days agoView details

  46. MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

    arXiv:2608.25085v1 Announce Type: new Abstract: Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a d…

    arxiv.org22 days agoView details

  47. Controllable Affective Generation via Latent Vector Steering

    arXiv:2608.25569v1 Announce Type: new Abstract: Large Language Models (LLMs) often produce emotionally flattened responses after alignment, limiting their effectiveness in affect-sensitive applications. In this paper, we propose EmoVec, a lightweight framework for controllable affective generation via latent vector st…

    arxiv.org22 days agoView details

  48. Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

    arXiv:2608.25574v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models…

    arxiv.org22 days agoView details

  49. Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification

    arXiv:2608.25579v1 Announce Type: new Abstract: Metaphor-identification performance can change markedly across datasets that differ in text distribution and annotation policy. We examine whether a fixed expert-informed procedure produces a more even cross-dataset profile than task-specific parameter adaptation. Four p…

    arxiv.org22 days agoView details

  50. AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

    arXiv:2608.25637v1 Announce Type: new Abstract: Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based, model-based, and tool-augmented verifie…

    arxiv.org22 days agoView details

  51. Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

    arXiv:2608.25869v1 Announce Type: new Abstract: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgment is often assumed to be independent…

    arxiv.org22 days agoView details

  52. GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding

    arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated correction pairs that are costly to maintain as query vocabularies evolve. Unsupervised correction with language models is…

    arxiv.org22 days agoView details

  53. MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

    arXiv:2608.25768v1 Announce Type: new Abstract: Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on a language-specifically filt…

    arxiv.org22 days agoView details

  54. Adaptive Triggering for Bias Correction in LLM Reasoning

    arXiv:2608.25379v1 Announce Type: new Abstract: Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation presents a fundamental timing problem: inte…

    arxiv.org22 days agoView details

  55. DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models

    arXiv:2608.25428v1 Announce Type: new Abstract: Correcting flawed reasoning traces remains a significant challenge for Large Language Models (LLMs), whose autoregressive generation can propagate early mistakes into subsequent reasoning. We introduce DCGC, a Masked Diffusion Model (MDM) framework for global correction…

    arxiv.org22 days agoView details

  56. Loss-Based Active Learning for Neural Abstractive Summarization

    arXiv:2608.25881v1 Announce Type: new Abstract: Fine-tuning abstractive summarization models requires high-quality annotated data. However, obtaining such corpora is expensive and time-consuming, as it requires human annotators to read and comprehend long documents to create accurate summaries. Active learning mitigat…

    arxiv.org22 days agoView details

  57. VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

    arXiv:2608.25478v1 Announce Type: new Abstract: In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI-generated text. It allows users to interact throug…

    arxiv.org22 days agoView details

  58. MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

    arXiv:2608.25449v1 Announce Type: new Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations…

    arxiv.org22 days agoView details

  59. Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens

    arXiv:2608.25347v1 Announce Type: new Abstract: The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal…

    arxiv.org22 days agoView details

  60. EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

    arXiv:2608.25561v1 Announce Type: new Abstract: VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between…

    arxiv.org22 days agoView details

  61. Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

    arXiv:2608.25654v1 Announce Type: new Abstract: Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs an…

    arxiv.org22 days agoView details

  62. The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

    arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It…

    arxiv.org22 days agoView details

  63. GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning

    arXiv:2608.25583v1 Announce Type: new Abstract: Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned models tend to answer more concisely, yet…

    arxiv.org22 days agoView details

  64. Localize-Then-Decide Guarantees for LLM Judgments

    arXiv:2608.25824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provid…

    arxiv.org22 days agoView details

  65. JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

    arXiv:2608.25593v1 Announce Type: new Abstract: Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual…

    arxiv.org22 days agoView details

  66. Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

    arXiv:2608.25761v1 Announce Type: new Abstract: One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the constrained case of this "model size vs.…

    arxiv.org22 days agoView details

  67. Virgil: Navigating Explainability for Transformer-based Language Models

    arXiv:2608.25555v1 Announce Type: new Abstract: Explainability for transformer-based language models is becoming crucial as these systems are deployed in high-stakes applications. As a result, the ecosystem of explainability tools is rapidly evolving, becoming richer, but also more fragmented and harder to navigate. T…

    arxiv.org22 days agoView details

  68. BanglaMamba: Exploring State Space Models for Bangla Fake News Detection

    arXiv:2608.25190v1 Announce Type: new Abstract: Fake news detection has become an important Natural Language Processing (NLP) task due to the rapid spread of misinformation through online news platforms and social media. While transformer-based models such as BanglaBERT achieve strong performance for Bangla text class…

    arxiv.org22 days agoView details

  69. One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

    arXiv:2608.25904v1 Announce Type: new Abstract: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), b…

    arxiv.org22 days agoView details

  70. When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies

    arXiv:2608.25717v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public…

    arxiv.org22 days agoView details

  71. SAMpLE: A SystemC-AMS Machine LEarning-based Framework for Virtual Prototyping

    arXiv:2608.25910v1 Announce Type: new Abstract: Machine Learning (ML) is increasingly used in virtual prototypes of embedded systems to model behaviors that are difficult to capture analytically. However, integrating ML models into virtual platform simulation is still typically done through ad hoc solutions, which lim…

    arxiv.org22 days agoView details

  72. OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

    arXiv:2608.25398v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we in…

    arxiv.org22 days agoView details

  73. Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

    arXiv:2608.25660v1 Announce Type: new Abstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their j…

    arxiv.org22 days agoView details

  74. Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

    arXiv:2608.25662v1 Announce Type: new Abstract: In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which…

    arxiv.org22 days agoView details

  75. The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers

    arXiv:2608.25166v1 Announce Type: new Abstract: Transformer representations describe trajectories through high-dimensional vector spaces, which are shaped dynamically as tokens incorporate relational context across layers. Such data tend to concentrate on lower-dimensional sub-manifolds, a form of compression quantifi…

    arxiv.org22 days agoView details

  76. Belief Cascades Drive Persuasion in LLM Agent Networks

    arXiv:2608.25152v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly debate answers, coordinate research, simulate users, and mediate information flows, making agent-to-agent persuasion a basic but undermeasured capability. We introduce a controlled testbed for studying how goal-directed persuaders shi…

    arxiv.org22 days agoView details

  77. From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations

    arXiv:2608.25894v1 Announce Type: new Abstract: Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturbations in the input. Existing approaches to hallucination mitigation typically assume reliab…

    arxiv.org22 days agoView details

  78. Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting

    arXiv:2608.25359v1 Announce Type: new Abstract: Conversational derailment forecasting aims to predict when online discussions will escalate into hostility, enabling proactive moderation. Existing approaches often struggle in low-data settings and to generalize across domains. This poses a challenge for new platforms a…

    arxiv.org22 days agoView details

  79. HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

    arXiv:2608.25071v1 Announce Type: new Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people t…

    arxiv.org22 days agoView details

  80. Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

    arXiv:2608.25826v1 Announce Type: new Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched. Scientific papers are wr…

    arxiv.org22 days agoView details

  81. Unsupervised Post-Training of Foundation Models: A Survey

    arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model a…

    arxiv.org22 days agoView details

  82. From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

    arXiv:2608.25605v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studies comparing prompted LLMs with state-of-the-art enc…

    arxiv.org22 days agoView details

  83. Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting

    arXiv:2608.25115v1 Announce Type: new Abstract: Existing methods for improving Retrieval-Augmented Generation (RAG) efficiency mainly optimize downstream LLM generation, such as context compression or serving optimization. However, RAG is an end-to-end system, and its bottleneck can shift between upstream reranking an…

    arxiv.org22 days agoView details

  84. Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

    arXiv:2608.25655v1 Announce Type: new Abstract: Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings u…

    arxiv.org22 days agoView details

  85. Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

    arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tu…

    arxiv.org22 days agoView details

  86. The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

    arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Usin…

    arxiv.org22 days agoView details

  87. Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark

    arXiv:2608.25854v1 Announce Type: new Abstract: Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating repre…

    arxiv.org22 days agoView details

  88. A Primer on Computational Semantics for Artificial Intelligence Systems

    arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is. This document is a…

    arxiv.org22 days agoView details

  89. TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

    arXiv:2608.25523v1 Announce Type: new Abstract: Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available for batching concurrent requests. In mul…

    arxiv.org22 days agoView details

  90. Padamitra: Grounded Glossary Generation for Classical Sanskrit

    arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluabl…

    arxiv.org22 days agoView details

  91. ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

    arXiv:2608.25531v1 Announce Type: new Abstract: Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long-context models. Compact, locally deployable language models are a…

    arxiv.org22 days agoView details

  92. Provenance Before Prose: Claim-Locked Reporting

    arXiv:2608.25336v1 Announce Type: new Abstract: Large language models (LLMs) can fluently verbalize statistical evidence, yet statistical reports can still drift numerical values, invert effect directions, or restate thresholded contrasts as categorical effects. We frame these failures as a control problem: the eviden…

    arxiv.org22 days agoView details

  93. Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

    arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and wit…

    arxiv.org22 days agoView details

  94. Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection

    arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning…

    arxiv.org22 days agoView details

  95. SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation

    arXiv:2608.25123v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language models by incorporating external knowledge without retraining, but existing methods often underuse the relational structure encoded in knowledge graphs. Graph-based RAG can capture entity relationships, yet sup…

    arxiv.org22 days agoView details

  1. zai-org/GLM-5.3-Flash

    image-text-to-text · transformers · safetensors · glm5_next

    huggingface.co22 days ago2408 ptsView details

  2. unsloth/GLM-5.3-Flash-GGUF

    text-generation · gguf · unsloth · glm5_next

    huggingface.co22 days ago376 ptsView details

  3. apodex/Apodex-1.1-mini

    text-generation · transformers · safetensors · qwen3_5_moe

    huggingface.co23 days ago107 ptsView details

  4. nlpai-lab/KURE-v1

    feature-extraction · sentence-transformers · safetensors · xlm-roberta

    huggingface.co23 days ago92 ptsView details

  5. AtomicChat/Qwen3.8-Flash-Next-GGUF

    text-generation · gguf · atomic-chat · qwen

    huggingface.co22 days ago82 ptsView details

  6. RadixArk/Qwen3.8-Flash-Next-NVFP4

    image-text-to-text · Model Optimizer · safetensors · qwen4_exp

    huggingface.co22 days ago79 ptsView details

  7. nlpai-lab/KoE5

    feature-extraction · transformers · safetensors · xlm-roberta

    huggingface.co23 days ago51 ptsView details

  8. quimmedes/Qwen3.8-27B-XYZ

    gguf · qwen · qwen3.5

    huggingface.co23 days ago40 ptsView details

  9. DZER-Studios/Vexion-gpt-medium

    text-generation · safetensors · vexion_gpt · pytorch

    huggingface.co23 days agoView details

  10. Ba2han/experimental4

    text-generation · transformers · safetensors · qwen3

    huggingface.co23 days agoView details

  11. sullivan1502/base-action-pretrain

    text-generation · transformers · safetensors · llama

    huggingface.co23 days agoView details

  12. kerasformers/glm-4.6v

    image-text-to-text · kerasformers · keras · glm

    huggingface.co23 days agoView details

  13. NotoriousH2/Qwen3-0.6B-JSON-SFT

    text-generation · safetensors · qwen3 · trl

    huggingface.co23 days agoView details

  14. kerasformers/glm-4.5-air-base

    text-generation · kerasformers · keras · glm

    huggingface.co23 days agoView details

  1. vllm-project/vllm v0.28.0

    # v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654,…

    github.com23 days ago108 ptsView details

  2. ggml-org/llama.cpp b10643

    <details open> hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchronous backend (#26501) * hexagon: use non-host bufs by default and make the backend fully async * hex-hb: remove optional hostbuf support and fix async copy * hex-unary: relax supported unary chec…

    github.com22 days agoView details

  3. openai/openai-python v3.5.0

    ## [3.5.0](https://github.com/openai/openai-python/compare/v3.4.0...v3.5.0) (2026-08-27) ### Features * **api:** make function call output call IDs optional ([#3738](https://github.com/openai/openai-python/issues/3738)) ([c74501d](https://github.com/openai/openai-python/commit/c…

    github.com22 days agoView details

  4. ggml-org/llama.cpp b10642

    <details open> llama: add token ID tracking to KV cell (#27762) * kv: track token id * rm get_prev_tokens, move it to the main pr * nits * add get_prev_tokens </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43…

    github.com22 days agoView details

  5. openai/openai-python v3.4.0

    ## [3.4.0](https://github.com/openai/openai-python/compare/v3.3.1...v3.4.0) (2026-08-25) ### Features * **api:** Add obfuscation field to ChatCompletionChunk ([#3690](https://github.com/openai/openai-python/issues/3690)) ([c7d8e1d](https://github.com/openai/openai-python/commit/…

    github.com22 days agoView details

  6. anthropics/anthropic-sdk-python v1.1.0

    ## 1.1.0 (2026-08-26) Full Changelog: [v1.0.0...v1.1.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.0.0...v1.1.0) ### Features * **api:** add `updates` thinking display mode (beta) ([eb4a73f](https://github.com/anthropics/anthropic-sdk-python/commit/eb4a73fdddc…

    github.com22 days agoView details

  7. huggingface/transformers v5.16.1

    # Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) # GLM-5.3-Flash <img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209" /> GLM-5.3-Flash, the first **natively multimo…

    github.com22 days agoView details

  8. huggingface/transformers v5.16.0

    # Release v5.16.0 ## New Model additions ### Qwen4-Exp <img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" /> Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key compone…

    github.com22 days agoView details

  9. ggml-org/llama.cpp b10631

    <details open> ggml-meta: propagate buffer usage and call init on the new tensors (#27586) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/43044002> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://githu…

    github.com23 days agoView details