Skip to content

Archive / 2026-08-31

August 31, 2026

  1. deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

    image-text-to-text · transformers · safetensors · deepseek_v4

    huggingface.co17 days ago874 ptsView details

  2. Agent memory as a file format

    calpaterson.com17 days ago191 ptsView detailsJoin discussion

  3. The safest job from AI may be writing

    muratbuffalo.blogspot.com17 days ago146 ptsView detailsJoin discussion

  4. Marx, Keynes, and AI

    unpopularfront.news17 days ago59 ptsView detailsJoin discussion

  5. AI-written code is still your code

    martiansoftware.com17 days ago58 ptsView detailsJoin discussion

  6. The British state has lost the argument

    jonathancook.substack.com17 days ago45 ptsView detailsJoin discussion

  7. ChatGPT to face tougher regulation in the EU

    OpenAI will soon be held accountable for mitigating risks related to ChatGPT's impact on minors, user mental health, and the spread of illegal content in the European Union. That's because ChatGPT is now considered a Very Large Online Search Engine under the EU's Digital Services Act, a set of laws regulating major on…

    theverge.com17 days ago25 ptsView detailsJoin discussion

  8. Agentic Trust Controls

    trustcontrols.ai17 days ago18 ptsView detailsJoin discussion

  9. How much of a problem is AI's water use?

    knowablemagazine.org17 days ago13 ptsView detailsJoin discussion

  10. You Know Who Really Hates AI? Insurance Claims Adjusters

    Of the Glassdoor reviews from claims adjusters that mentioned AI, a staggering 98 percent were negative. “AI is just a tool,” one person tells WIRED. “It should never be given the keys.”

    wired.com17 days ago11 ptsView details

  11. "Dagon" by H. P. Lovecraft

    hplovecraft.com17 days ago10 ptsView detailsJoin discussion

  12. Weird Little Guys of FOSS

    drewdevault.com17 days ago10 ptsView detailsJoin discussion

  13. Primary source

    How law firm Gilbert + Tobin governs and scales AI with OpenAI

    See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.

    openai.com17 days agoView details

  14. Primary source

    OpenAI supports California’s bill to advance youth AI safety

    OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.

    openai.com18 days agoView details

  15. Primary source

    Polimill builds Japan's next-generation public AI infrastructure

    Polimill uses OpenAI GPT models and Codex to help municipalities search and use administrative knowledge while accelerating development.

    openai.com18 days agoView details

  16. Primary source

    TimesFM-3: A zero-shot foundation model for multivariate forecasting

    Data Management

    research.google17 days agoView details

  17. Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio

    Speed and accuracy usually pull against each other in text-to-speech. Gradium AI's new default model reports both: an 81.0% human-rated pass rate on 500 hard sentences across five languages, at 216 ms P50 time-to-first-audio on Coval. The evaluation set is open on Hugging Face under CC BY 4.0. The post Gradium AI Rele…

    marktechpost.com17 days agoView details

  18. Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour

    How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in […] The post K…

    marktechpost.com17 days agoView details

  19. Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting

    Google Research has released TimesFM-3, a 330 million parameter time series foundation model that forecasts multiple related series in a single forward pass. Unlike every TimesFM checkpoint through 2.5, it is pretrained natively for multivariate forecasting, accepting multiple targets, past covariates, and past-future…

    marktechpost.com17 days agoView details

  20. OpenClaw Releases OpenClaw 2.0: Guided Model Setup, 575 ms Control UI Startup, and One Trust Boundary Per Gateway

    The OpenClaw Foundation has released v2026.8.1, which the project calls OpenClaw 2.0: 933 contributors, 569 first-timers, and more than 16,000 pull requests, roughly half of every PR ever merged into the repo. Setup now reuses existing subscriptions, API keys and local models. The rebuilt Control UI cut test-harness s…

    marktechpost.com18 days agoView details

  21. Debian won’t ban AI code from its Linux distribution

    Debian voted to allow developers to use AI tools in their contributions to the Linux distribution's "development, maintenance, [and] documentation." The new policy on AI acknowledges that "responsible" use of AI can improve developers' productivity, and goes on to say, "generative AI is neither exempt from nor subject…

    theverge.com17 days agoView details

  22. New York Governor Kathy Hochul thinks AI should be ‘less evil’

    Today, I’m talking with New York Governor Kathy Hochul, and I’ll just warn you — this episode moves really fast. It’s an election year, after all, with a shocking amount of tech policy at stake, and Governor Hochul has taken strong positions on almost every major tech issue there is. For example, Meta just reached a s…

    theverge.com17 days agoView details

  23. Instagram cracks down on AI accounts pretending to be human

    “AI creator” accounts like Aitana Lopez will get a new “AI-generated profile” label. | Image: Aitana Lopez Instagram is finally taking steps to address the rise of fake AI-influencer accounts that have gotten harder to spot. It's also renaming the "AI creator" label to "AI-generated profile" to make it clear when a pr…

    theverge.com17 days agoView details

  24. Parametric Multimodal User Memory: Storing What Captions Cannot Carry

    arXiv:2608.28609v1 Announce Type: new Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved by similarity. This serves the captionable half of a person ("my cat is named Bibi"), but discards the perceptual half no…

    arxiv.org17 days agoView details

  25. Leveraging Turn-taking Dynamics for Intent Recognition in Multi-party Conversations

    arXiv:2608.28926v1 Announce Type: new Abstract: We propose a multi-task learning approach for multi-party dialogue intent recognition that leverages an auxiliary task that models turn-taking dynamics. Specifically, we introduce turn-transition entropy, a self-supervised target computed from the sequence of speaker tra…

    arxiv.org17 days agoView details

  26. StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue

    arXiv:2608.29326v1 Announce Type: new Abstract: Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic replies but also coherent progression through a multi-turn support process. Existing resources often reduce supervision to turn…

    arxiv.org17 days agoView details

  27. Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

    arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a mod…

    arxiv.org17 days agoView details

  28. MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments

    arXiv:2608.28624v1 Announce Type: new Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-s…

    arxiv.org17 days agoView details

  29. PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

    arXiv:2608.28640v1 Announce Type: new Abstract: In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to ef…

    arxiv.org17 days agoView details

  30. Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

    arXiv:2608.28641v1 Announce Type: new Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Ko…

    arxiv.org17 days agoView details

  31. Redesigning and Auditing Deep Research Writing for Faithful Reports

    arXiv:2608.28643v1 Announce Type: new Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and…

    arxiv.org17 days agoView details

  32. GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon

    arXiv:2608.28667v1 Announce Type: new Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple Silicon, with its un…

    arxiv.org17 days agoView details

  33. ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

    arXiv:2608.28707v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as obje…

    arxiv.org17 days agoView details

  34. No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

    arXiv:2608.28875v1 Announce Type: new Abstract: Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the prompt-conditioning l…

    arxiv.org17 days agoView details

  35. VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition

    arXiv:2608.28916v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values for identifiers, paths, and measured quantities. A transcript can appear fluent and achieve low WER while corrupting a va…

    arxiv.org17 days agoView details

  36. CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

    arXiv:2608.28958v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkwa…

    arxiv.org17 days agoView details

  37. The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era

    arXiv:2608.28980v1 Announce Type: new Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This question is examined through a review of 159 papers (2016--2026) across nine modalities, with predictive accuracy considered alo…

    arxiv.org17 days agoView details

  38. Quantifying Error Tolerance in Synthetic Data: An Atomic-level Operand vs. Operator Perturbation Study

    arXiv:2608.29144v1 Announce Type: new Abstract: Synthetic data generation has become a cornerstone for advancing large language models. However, the lack of the quantitative analysis for error tolerance became a critical bottleneck. Consequently, current filtering strategies fluctuate between two extremes: they are ei…

    arxiv.org17 days agoView details

  39. STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study

    arXiv:2608.28614v1 Announce Type: new Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies…

    arxiv.org17 days agoView details

  40. AI Historian: Helping historians organize and verify person-centred temporal clues from dispersed historical narratives

    arXiv:2608.29133v1 Announce Type: new Abstract: History is not preserved in complete, continuous form. Accounts of a person's activities, relationships and historical contexts are scattered across texts, chapters and narrative perspectives; historians must retrieve, identify and compare these materials to reconstruct…

    arxiv.org17 days agoView details

  41. Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities

    arXiv:2608.29209v1 Announce Type: new Abstract: In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture represe…

    arxiv.org17 days agoView details

  42. When Patients Cut In: Extending Clinical Conversational AI Safety to Interruptions

    arXiv:2608.29241v1 Announce Type: new Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a cascaded architecture (speech-to-text -> LLM -> text-to-speech), so when a patient cuts the agent off mid-utterance, clinical…

    arxiv.org17 days agoView details

  43. Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict

    arXiv:2608.28645v1 Announce Type: new Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments…

    arxiv.org17 days agoView details

  44. Not All or None: Dynamic Construction of Target-aware Memory Graph for Conversational Stance Detection

    arXiv:2608.29066v1 Announce Type: new Abstract: Stance detection is crucial for understanding the underlying attitude of an expression towards a target. Conversational stance detection is a more challenging stance detection task in real-world social media scenarios, as it involves detecting the user's stance by levera…

    arxiv.org17 days agoView details

  45. Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning

    arXiv:2608.29284v1 Announce Type: new Abstract: Deploying large language models for legal question answering raises challenges that general-purpose leaderboards do not capture, particularly for low-resource languages and under hard operational constraints. We report on building and operating a retrieval-augmented (RAG…

    arxiv.org17 days agoView details

  46. Learning Simple Test-Time Environments for LLM Web Agents

    arXiv:2608.29305v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real-world settings. Existing research largely attribute this degradation to the compo…

    arxiv.org17 days agoView details

  47. Pad\=artha: Ontology-Grounded Fine-Grained NER Benchmark for Classical Sanskrit

    arXiv:2608.29324v1 Announce Type: new Abstract: Annotation schemas are not neutral. When applied to classical literature, tag sets developed for modern journalistic texts impose source-culture definitions on texts they were never designed to describe. We instead ground a schema in the tradition of the text itself intr…

    arxiv.org17 days agoView details

  48. AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

    arXiv:2608.29397v1 Announce Type: new Abstract: Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real-world decision settings such as route planning and fleet dispatch. Individual choices interact thr…

    arxiv.org17 days agoView details

  49. Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?

    arXiv:2608.28649v1 Announce Type: new Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on collaborative-filtering-based heuristics, w…

    arxiv.org17 days agoView details

  50. Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

    arXiv:2608.29446v1 Announce Type: new Abstract: Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single,…

    arxiv.org17 days agoView details

  51. Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration

    arXiv:2608.29378v1 Announce Type: new Abstract: Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal rather than harmful co…

    arxiv.org17 days agoView details

  52. Test-Time Scaling for Scientific Equation Discovery

    arXiv:2608.28660v1 Announce Type: new Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an open-ended setting where models search over c…

    arxiv.org17 days agoView details

  53. VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models

    arXiv:2608.28932v1 Announce Type: new Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. The benchmark…

    arxiv.org17 days agoView details

  54. Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study

    arXiv:2608.29170v1 Announce Type: new Abstract: This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure and a community-driven open-source workflow. The design framework addresses the lack of systematic cross-lingual romanizat…

    arxiv.org17 days agoView details

  55. Latent-Space Intervention for Cross-Lingual Factual Consistency: Consistency Improvements without Accuracy Drops

    arXiv:2608.28860v1 Announce Type: new Abstract: Large Language Models (LLMs) often answer the same factual question differently across languages. We study whether cross-lingual latent-space intervention can reduce this inconsistency. We train layer-specific autoencoders on parallel multilingual representations and app…

    arxiv.org17 days agoView details

  56. A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

    arXiv:2608.29034v1 Announce Type: new Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language mode…

    arxiv.org17 days agoView details

  57. SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization

    arXiv:2608.29270v1 Announce Type: new Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in a different formulati…

    arxiv.org17 days agoView details

  58. Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

    arXiv:2608.29278v1 Announce Type: new Abstract: Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tel…

    arxiv.org17 days agoView details

  59. Evaluating the Semantic Specificity of Representation Steering in Language Models

    arXiv:2608.29431v1 Announce Type: new Abstract: Localized Representation Steering (LRS) is widely used to correct reasoning pathologies in large language models. However, standard benchmark evaluations can easily be fooled by superficial label overrides, creating a false impression of reasoning circuit repairs. In thi…

    arxiv.org17 days agoView details

  60. Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

    arXiv:2608.28630v1 Announce Type: new Abstract: Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This creates a key challen…

    arxiv.org17 days agoView details

  61. Detecting and Guiding LLM-Generated Korean Poetry with Interpretable Form-level Features

    arXiv:2608.28986v1 Announce Type: new Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a Korean poem is human- or LLM-authored, and guiding LLMs to generate poetry closer in form to human writing. We quantify t…

    arxiv.org17 days agoView details

  62. Detecting and Repairing Hallucinations in Retrieval-Augmented Generation

    arXiv:2608.29307v1 Announce Type: new Abstract: Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants and enterprise knowledge tools. Grounding a model in retrieved text reduces unsupported statements but does not eliminate th…

    arxiv.org17 days agoView details

  63. When to Adapt: Conditional Memory Adapters for Retention-Preserving Domain Specialization

    arXiv:2608.29327v1 Announce Type: new Abstract: Large language models deployed in specialized domains must improve in-domain performance without sacrificing general capabilities. Existing parameter-efficient fine-tuning methods are typically always on: their learned perturbations are applied to every input, which can…

    arxiv.org17 days agoView details

  64. Asymmetric Within-Document Predictive Learning for Scientific Document Representation

    arXiv:2608.28625v1 Announce Type: new Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict…

    arxiv.org17 days agoView details

  65. Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

    arXiv:2608.28629v1 Announce Type: new Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM. Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs. Firstly, a BIM-to-Text…

    arxiv.org17 days agoView details

  66. All You Need Is Non-Commutative Words

    arXiv:2608.29314v1 Announce Type: new Abstract: We represent lexical tokens as unitary matrices and encode each sentence as their ordered product. The noncommutativity of matrix product captures word order without positional encodings (PEs). The same algebra yields several capabilities, including antisymmetric self-at…

    arxiv.org17 days agoView details

  67. PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation

    arXiv:2608.28633v1 Announce Type: new Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation st…

    arxiv.org17 days agoView details

  68. Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

    arXiv:2608.28635v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges beca…

    arxiv.org17 days agoView details

  69. From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

    arXiv:2608.28619v1 Announce Type: new Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable…

    arxiv.org17 days agoView details

  70. Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

    arXiv:2608.28776v1 Announce Type: new Abstract: Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multilingual embedding mod…

    arxiv.org17 days agoView details

  71. HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding

    arXiv:2608.29120v1 Announce Type: new Abstract: Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnos…

    arxiv.org17 days agoView details

  72. A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

    arXiv:2608.28846v1 Announce Type: new Abstract: Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time…

    arxiv.org17 days agoView details

  73. Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

    arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026…

    arxiv.org17 days agoView details

  74. Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

    arXiv:2608.28611v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian education system is li…

    arxiv.org17 days agoView details

  75. Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

    arXiv:2608.29215v1 Announce Type: new Abstract: To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. So far, prompting alone has been shown to be insufficient for creating such explanations and other computational methods are missing. Therefore, thi…

    arxiv.org17 days agoView details

  76. Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions

    arXiv:2608.29109v1 Announce Type: new Abstract: Large language models often answer structurally unanswerable questions, such as computing cot(-540{\deg}) or evaluating (1).startswith("1"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention.…

    arxiv.org17 days agoView details

  77. Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

    arXiv:2608.29239v1 Announce Type: new Abstract: Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translat…

    arxiv.org17 days agoView details

  78. Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs

    arXiv:2608.29257v1 Announce Type: new Abstract: Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over less popular correct ones, a vulnerabili…

    arxiv.org17 days agoView details

  79. Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models

    arXiv:2608.28924v1 Announce Type: new Abstract: Linguistic theory has long recognized cross-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar mechanisms. However, this hypothesis has been difficult to test empirically due to our lack of fine-grained, manipulabl…

    arxiv.org17 days agoView details

  80. NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

    arXiv:2608.28608v1 Announce Type: new Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The research here utilize…

    arxiv.org17 days agoView details

  81. Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

    arXiv:2608.28886v1 Announce Type: new Abstract: What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 poin…

    arxiv.org17 days agoView details

  82. RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling

    arXiv:2608.29058v1 Announce Type: new Abstract: Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and estimates that pattern's sparse indices for every prompt. This design is efficient, but it ass…

    arxiv.org17 days agoView details

  83. The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

    arXiv:2608.28930v1 Announce Type: new Abstract: Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal over…

    arxiv.org17 days agoView details

  1. google/timesfm-3.0-pytorch

    time-series-forecasting · safetensors · time-series · forecasting

    huggingface.co17 days ago832 ptsView details

  2. OpenMOSS-Team/MOSS-Transcribe-Diarize

    audio-text-to-text · transformers · safetensors · moss_transcribe_diarize

    huggingface.co18 days ago406 ptsView details

  3. lightonai/ColBERT-Zero

    sentence-similarity · PyLate · safetensors · modernbert

    huggingface.co18 days ago44 ptsView details

  4. litert-community/Qwen3.5-0.8B

    text-generation · litert-lm · litert · litertlm

    huggingface.co18 days ago3 ptsView details

  5. VaultLevel6/beaverai_gguf_backups

    gguf · endpoints_compatible · region:us

    huggingface.co18 days agoView details

  6. mradermacher/Puro-2B-Base-i1-GGUF

    transformers · gguf · base-model

    huggingface.co18 days agoView details

  7. ishikaa/acquisition_student_AS_format_numina_qwen7b

    text-generation · transformers · safetensors · qwen2

    huggingface.co18 days agoView details

  8. VertexAGI/amethyst-1-mini

    text-generation · transformers · safetensors · gguf

    huggingface.co18 days agoView details

  9. VertexAGI/prism-roleplay-1.5-small

    text-generation · mlx · safetensors · gguf

    huggingface.co18 days agoView details

  10. VertexAGI/prism-roleplay-1-small

    text-generation · mlx · safetensors · gguf

    huggingface.co18 days agoView details

  1. ggml-org/llama.cpp b10729

    <details open> metal : add fa-vec tunings for M1 Ultra (#28088) * metal : add fa-vec tunings for M1 Ultra * metal : move M1 Ultra tunings after M1 Max section * metal : remove duplicate blank line </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.…

    github.com17 days agoView details

  2. modelcontextprotocol/servers 2026.8.31

    # Release : v2026.8.31 ## Updated packages - @modelcontextprotocol/server-filesystem@2026.8.31 - @modelcontextprotocol/server-memory@2026.8.31 - @modelcontextprotocol/server-sequential-thinking@2026.8.31 - @modelcontextprotocol/server-everything@2026.8.31

    github.com17 days agoView details

  3. ggml-org/llama.cpp b10712

    <details open> vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (#28032) * vulkan: add top-k radix sort shader for k >= 1024 * add Qwen 3.8 Flash Next top-k tests * add top-k qsa fusion * clean up code </details> **Website:** - <https://llama.app> **Attestations:…

    github.com18 days agoView details