Archive / 2026-08-31
August 31, 2026
News
View all news →deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
image-text-to-text · transformers · safetensors · deepseek_v4
huggingface.co17 days ago874 ptsView details
Apple caught off guard by AI demand for Mac Mini and Mac Studio
macrumors.com17 days ago492 ptsView detailsJoin discussion
RavynOS: Pre-alpha open-source OS based on Darwin, FreeBSD, Apple open-source
ravynos.com17 days ago233 ptsView detailsJoin discussion
calpaterson.com17 days ago191 ptsView detailsJoin discussion
The safest job from AI may be writing
muratbuffalo.blogspot.com17 days ago146 ptsView detailsJoin discussion
2004 RuneScape fit a multiplayer RPG into 56k dial-up
jkm.dev17 days ago138 ptsView detailsJoin discussion
Rakuten Kobo returns to U.S. retail as sales double
publishersweekly.com17 days ago105 ptsView detailsJoin discussion
Ex-Crips leader found guilty in 1996 murder of rapper Tupac Shakur
bbc.com17 days ago94 ptsView detailsJoin discussion
DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
dolthub.com17 days ago62 ptsView detailsJoin discussion
Claude Code reduces it's weekly limit by 17% – compared to today
twitter.com18 days ago67 ptsView detailsJoin discussion
Launch HN: Almanac (YC S26) – AI that knows your company
usealmanac.com17 days ago59 ptsView detailsJoin discussion
unpopularfront.news17 days ago59 ptsView detailsJoin discussion
Meta Security Researcher's AI Agent Accidentally Deleted Her Emails
au.pcmag.com18 days ago60 ptsView detailsJoin discussion
AI-written code is still your code
martiansoftware.com17 days ago58 ptsView detailsJoin discussion
Show HN: Corporate Mind Games – logic puzzles with a sarcastic corporate theme
corporatemindgames.com17 days ago49 ptsView detailsJoin discussion
Launch HN: Hebbian Robotics (YC S26) – Build scalable robotics data pipelines
github.com17 days ago46 ptsView detailsJoin discussion
DIY archivists push budget Nikons to 902,000 clicks to save 1,800 rare books
tomshardware.com17 days ago46 ptsView detailsJoin discussion
The British state has lost the argument
jonathancook.substack.com17 days ago45 ptsView detailsJoin discussion
Apple Is Suddenly an AI Infra Stock as OpenAI Buys 10k+ Macs
247wallst.com17 days ago39 ptsView detailsJoin discussion
California Lawmakers Pass Plug-In Solar Bill
nytimes.com17 days ago34 ptsView detailsJoin discussion
ArXiv has almost 600 submissions today and most are AI slop
arxiv.org17 days ago30 ptsView detailsJoin discussion
How we configured OpenTelemetry logs in Rails
sixpatterns.com17 days ago30 ptsView detailsJoin discussion
ChatGPT to face tougher regulation in the EU
OpenAI will soon be held accountable for mitigating risks related to ChatGPT's impact on minors, user mental health, and the spread of illegal content in the European Union. That's because ChatGPT is now considered a Very Large Online Search Engine under the EU's Digital Services Act, a set of laws regulating major on…
theverge.com17 days ago25 ptsView detailsJoin discussion
Wildfire Survivors Push Back Against Newsom-Led Utility Bailout
prospect.org17 days ago25 ptsView detailsJoin discussion
What I Learned About AI Trust from Reconciling over 100B Transactions
engineering.moniepoint.com17 days ago25 ptsView detailsJoin discussion
Show HN: SlideOps – slides from a repo that flag when they drift from the code
github.com17 days ago23 ptsView detailsJoin discussion
Show HN: Floe – an open-source plugin for sample libraries – CLAP/VST3/AU
floe.audio18 days ago22 ptsView detailsJoin discussion
Show HN: 49 IDE – 2D Canvas for Agents
github.com17 days ago18 ptsView detailsJoin discussion
trustcontrols.ai17 days ago18 ptsView detailsJoin discussion
How much of a problem is AI's water use?
knowablemagazine.org17 days ago13 ptsView detailsJoin discussion
New US missile left wide path of death and destruction in Iranian neighborhoods
apnews.com17 days ago13 ptsView detailsJoin discussion
Lambs grazing under an Oregon solar farm put on the same weight to the gram
spacedaily.com18 days ago13 ptsView detailsJoin discussion
You Know Who Really Hates AI? Insurance Claims Adjusters
Of the Glassdoor reviews from claims adjusters that mentioned AI, a staggering 98 percent were negative. “AI is just a tool,” one person tells WIRED. “It should never be given the keys.”
wired.com17 days ago11 ptsView details
Run multiple Linux kernels without hypervisor
lore.kernel.org18 days ago12 ptsView detailsJoin discussion
EU Commission Designates ChatGPT, Reddit and Roblox Under the DSA
digital-strategy.ec.europa.eu17 days ago11 ptsView detailsJoin discussion
Sabine Hossenfelder: More Exposed Than Ever (Racism Edition) [video]
youtube.com17 days ago11 ptsView detailsJoin discussion
The AI-Native SDLC Starts with Your Infrastructure
metalbear.com17 days ago11 ptsView detailsJoin discussion
hplovecraft.com17 days ago10 ptsView detailsJoin discussion
Wendell Berry has died, age 92
nytimes.com17 days ago10 ptsView detailsJoin discussion
drewdevault.com17 days ago10 ptsView detailsJoin discussion
AI ruined some of my most precious accessibility tools
straye.dgirl.gay17 days ago10 ptsView detailsJoin discussion
Show HN: What Happens When You Give Your AI Agents a Voice and an Attitude
fellowgeek.github.io17 days ago10 ptsView detailsJoin discussion
- Primary source
How law firm Gilbert + Tobin governs and scales AI with OpenAI
See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.
openai.com17 days agoView details
- Primary source
OpenAI supports California’s bill to advance youth AI safety
OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.
openai.com18 days agoView details
- Primary source
Polimill builds Japan's next-generation public AI infrastructure
Polimill uses OpenAI GPT models and Codex to help municipalities search and use administrative knowledge while accelerating development.
openai.com18 days agoView details
- Primary source
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
huggingface.co17 days agoView details
- Primary source
TimesFM-3: A zero-shot foundation model for multivariate forecasting
Data Management
research.google17 days agoView details
Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio
Speed and accuracy usually pull against each other in text-to-speech. Gradium AI's new default model reports both: an 81.0% human-rated pass rate on 500 hard sentences across five languages, at 216 ms P50 time-to-first-audio on Coval. The evaluation set is open on Hugging Face under CC BY 4.0. The post Gradium AI Rele…
marktechpost.com17 days agoView details
Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour
How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in […] The post K…
marktechpost.com17 days agoView details
Google Research has released TimesFM-3, a 330 million parameter time series foundation model that forecasts multiple related series in a single forward pass. Unlike every TimesFM checkpoint through 2.5, it is pretrained natively for multivariate forecasting, accepting multiple targets, past covariates, and past-future…
marktechpost.com17 days agoView details
The OpenClaw Foundation has released v2026.8.1, which the project calls OpenClaw 2.0: 933 contributors, 569 first-timers, and more than 16,000 pull requests, roughly half of every PR ever merged into the repo. Setup now reuses existing subscriptions, API keys and local models. The rebuilt Control UI cut test-harness s…
marktechpost.com18 days agoView details
Debian won’t ban AI code from its Linux distribution
Debian voted to allow developers to use AI tools in their contributions to the Linux distribution's "development, maintenance, [and] documentation." The new policy on AI acknowledges that "responsible" use of AI can improve developers' productivity, and goes on to say, "generative AI is neither exempt from nor subject…
theverge.com17 days agoView details
New York Governor Kathy Hochul thinks AI should be ‘less evil’
Today, I’m talking with New York Governor Kathy Hochul, and I’ll just warn you — this episode moves really fast. It’s an election year, after all, with a shocking amount of tech policy at stake, and Governor Hochul has taken strong positions on almost every major tech issue there is. For example, Meta just reached a s…
theverge.com17 days agoView details
Instagram cracks down on AI accounts pretending to be human
“AI creator” accounts like Aitana Lopez will get a new “AI-generated profile” label. | Image: Aitana Lopez Instagram is finally taking steps to address the rise of fake AI-influencer accounts that have gotten harder to spot. It's also renaming the "AI creator" label to "AI-generated profile" to make it clear when a pr…
theverge.com17 days agoView details
Parametric Multimodal User Memory: Storing What Captions Cannot Carry
arXiv:2608.28609v1 Announce Type: new Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved by similarity. This serves the captionable half of a person ("my cat is named Bibi"), but discards the perceptual half no…
arxiv.org17 days agoView details
Leveraging Turn-taking Dynamics for Intent Recognition in Multi-party Conversations
arXiv:2608.28926v1 Announce Type: new Abstract: We propose a multi-task learning approach for multi-party dialogue intent recognition that leverages an auxiliary task that models turn-taking dynamics. Specifically, we introduce turn-transition entropy, a self-supervised target computed from the sequence of speaker tra…
arxiv.org17 days agoView details
StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue
arXiv:2608.29326v1 Announce Type: new Abstract: Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic replies but also coherent progression through a multi-turn support process. Existing resources often reduce supervision to turn…
arxiv.org17 days agoView details
Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure
arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a mod…
arxiv.org17 days agoView details
arXiv:2608.28624v1 Announce Type: new Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-s…
arxiv.org17 days agoView details
PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework
arXiv:2608.28640v1 Announce Type: new Abstract: In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to ef…
arxiv.org17 days agoView details
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
arXiv:2608.28641v1 Announce Type: new Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Ko…
arxiv.org17 days agoView details
Redesigning and Auditing Deep Research Writing for Faithful Reports
arXiv:2608.28643v1 Announce Type: new Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and…
arxiv.org17 days agoView details
arXiv:2608.28667v1 Announce Type: new Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple Silicon, with its un…
arxiv.org17 days agoView details
ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering
arXiv:2608.28707v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as obje…
arxiv.org17 days agoView details
arXiv:2608.28875v1 Announce Type: new Abstract: Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the prompt-conditioning l…
arxiv.org17 days agoView details
VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
arXiv:2608.28916v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values for identifiers, paths, and measured quantities. A transcript can appear fluent and achieve low WER while corrupting a va…
arxiv.org17 days agoView details
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
arXiv:2608.28958v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkwa…
arxiv.org17 days agoView details
arXiv:2608.28980v1 Announce Type: new Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This question is examined through a review of 159 papers (2016--2026) across nine modalities, with predictive accuracy considered alo…
arxiv.org17 days agoView details
arXiv:2608.29144v1 Announce Type: new Abstract: Synthetic data generation has become a cornerstone for advancing large language models. However, the lack of the quantitative analysis for error tolerance became a critical bottleneck. Consequently, current filtering strategies fluctuate between two extremes: they are ei…
arxiv.org17 days agoView details
STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study
arXiv:2608.28614v1 Announce Type: new Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies…
arxiv.org17 days agoView details
arXiv:2608.29133v1 Announce Type: new Abstract: History is not preserved in complete, continuous form. Accounts of a person's activities, relationships and historical contexts are scattered across texts, chapters and narrative perspectives; historians must retrieve, identify and compare these materials to reconstruct…
arxiv.org17 days agoView details
arXiv:2608.29209v1 Announce Type: new Abstract: In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture represe…
arxiv.org17 days agoView details
When Patients Cut In: Extending Clinical Conversational AI Safety to Interruptions
arXiv:2608.29241v1 Announce Type: new Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a cascaded architecture (speech-to-text -> LLM -> text-to-speech), so when a patient cuts the agent off mid-utterance, clinical…
arxiv.org17 days agoView details
arXiv:2608.28645v1 Announce Type: new Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments…
arxiv.org17 days agoView details
arXiv:2608.29066v1 Announce Type: new Abstract: Stance detection is crucial for understanding the underlying attitude of an expression towards a target. Conversational stance detection is a more challenging stance detection task in real-world social media scenarios, as it involves detecting the user's stance by levera…
arxiv.org17 days agoView details
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning
arXiv:2608.29284v1 Announce Type: new Abstract: Deploying large language models for legal question answering raises challenges that general-purpose leaderboards do not capture, particularly for low-resource languages and under hard operational constraints. We report on building and operating a retrieval-augmented (RAG…
arxiv.org17 days agoView details
Learning Simple Test-Time Environments for LLM Web Agents
arXiv:2608.29305v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real-world settings. Existing research largely attribute this degradation to the compo…
arxiv.org17 days agoView details
Pad\=artha: Ontology-Grounded Fine-Grained NER Benchmark for Classical Sanskrit
arXiv:2608.29324v1 Announce Type: new Abstract: Annotation schemas are not neutral. When applied to classical literature, tag sets developed for modern journalistic texts impose source-culture definitions on texts they were never designed to describe. We instead ground a schema in the tradition of the text itself intr…
arxiv.org17 days agoView details
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
arXiv:2608.29397v1 Announce Type: new Abstract: Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real-world decision settings such as route planning and fleet dispatch. Individual choices interact thr…
arxiv.org17 days agoView details
Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?
arXiv:2608.28649v1 Announce Type: new Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on collaborative-filtering-based heuristics, w…
arxiv.org17 days agoView details
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts
arXiv:2608.29446v1 Announce Type: new Abstract: Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single,…
arxiv.org17 days agoView details
Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration
arXiv:2608.29378v1 Announce Type: new Abstract: Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal rather than harmful co…
arxiv.org17 days agoView details
Test-Time Scaling for Scientific Equation Discovery
arXiv:2608.28660v1 Announce Type: new Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an open-ended setting where models search over c…
arxiv.org17 days agoView details
VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models
arXiv:2608.28932v1 Announce Type: new Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. The benchmark…
arxiv.org17 days agoView details
arXiv:2608.29170v1 Announce Type: new Abstract: This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure and a community-driven open-source workflow. The design framework addresses the lack of systematic cross-lingual romanizat…
arxiv.org17 days agoView details
arXiv:2608.28860v1 Announce Type: new Abstract: Large Language Models (LLMs) often answer the same factual question differently across languages. We study whether cross-lingual latent-space intervention can reduce this inconsistency. We train layer-specific autoencoders on parallel multilingual representations and app…
arxiv.org17 days agoView details
arXiv:2608.29034v1 Announce Type: new Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language mode…
arxiv.org17 days agoView details
SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization
arXiv:2608.29270v1 Announce Type: new Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in a different formulati…
arxiv.org17 days agoView details
Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning
arXiv:2608.29278v1 Announce Type: new Abstract: Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tel…
arxiv.org17 days agoView details
Evaluating the Semantic Specificity of Representation Steering in Language Models
arXiv:2608.29431v1 Announce Type: new Abstract: Localized Representation Steering (LRS) is widely used to correct reasoning pathologies in large language models. However, standard benchmark evaluations can easily be fooled by superficial label overrides, creating a false impression of reasoning circuit repairs. In thi…
arxiv.org17 days agoView details
Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
arXiv:2608.28630v1 Announce Type: new Abstract: Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This creates a key challen…
arxiv.org17 days agoView details
Detecting and Guiding LLM-Generated Korean Poetry with Interpretable Form-level Features
arXiv:2608.28986v1 Announce Type: new Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a Korean poem is human- or LLM-authored, and guiding LLMs to generate poetry closer in form to human writing. We quantify t…
arxiv.org17 days agoView details
Detecting and Repairing Hallucinations in Retrieval-Augmented Generation
arXiv:2608.29307v1 Announce Type: new Abstract: Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants and enterprise knowledge tools. Grounding a model in retrieved text reduces unsupported statements but does not eliminate th…
arxiv.org17 days agoView details
When to Adapt: Conditional Memory Adapters for Retention-Preserving Domain Specialization
arXiv:2608.29327v1 Announce Type: new Abstract: Large language models deployed in specialized domains must improve in-domain performance without sacrificing general capabilities. Existing parameter-efficient fine-tuning methods are typically always on: their learned perturbations are applied to every input, which can…
arxiv.org17 days agoView details
Asymmetric Within-Document Predictive Learning for Scientific Document Representation
arXiv:2608.28625v1 Announce Type: new Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict…
arxiv.org17 days agoView details
arXiv:2608.28629v1 Announce Type: new Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM. Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs. Firstly, a BIM-to-Text…
arxiv.org17 days agoView details
All You Need Is Non-Commutative Words
arXiv:2608.29314v1 Announce Type: new Abstract: We represent lexical tokens as unitary matrices and encode each sentence as their ordered product. The noncommutativity of matrix product captures word order without positional encodings (PEs). The same algebra yields several capabilities, including antisymmetric self-at…
arxiv.org17 days agoView details
PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation
arXiv:2608.28633v1 Announce Type: new Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation st…
arxiv.org17 days agoView details
Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA
arXiv:2608.28635v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges beca…
arxiv.org17 days agoView details
arXiv:2608.28619v1 Announce Type: new Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable…
arxiv.org17 days agoView details
arXiv:2608.28776v1 Announce Type: new Abstract: Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multilingual embedding mod…
arxiv.org17 days agoView details
HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding
arXiv:2608.29120v1 Announce Type: new Abstract: Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnos…
arxiv.org17 days agoView details
arXiv:2608.28846v1 Announce Type: new Abstract: Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time…
arxiv.org17 days agoView details
arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026…
arxiv.org17 days agoView details
Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System
arXiv:2608.28611v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian education system is li…
arxiv.org17 days agoView details
Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation
arXiv:2608.29215v1 Announce Type: new Abstract: To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. So far, prompting alone has been shown to be insufficient for creating such explanations and other computational methods are missing. Therefore, thi…
arxiv.org17 days agoView details
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
arXiv:2608.29109v1 Announce Type: new Abstract: Large language models often answer structurally unanswerable questions, such as computing cot(-540{\deg}) or evaluating (1).startswith("1"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention.…
arxiv.org17 days agoView details
arXiv:2608.29239v1 Announce Type: new Abstract: Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translat…
arxiv.org17 days agoView details
Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs
arXiv:2608.29257v1 Announce Type: new Abstract: Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over less popular correct ones, a vulnerabili…
arxiv.org17 days agoView details
arXiv:2608.28924v1 Announce Type: new Abstract: Linguistic theory has long recognized cross-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar mechanisms. However, this hypothesis has been difficult to test empirically due to our lack of fine-grained, manipulabl…
arxiv.org17 days agoView details
arXiv:2608.28608v1 Announce Type: new Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The research here utilize…
arxiv.org17 days agoView details
arXiv:2608.28886v1 Announce Type: new Abstract: What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 poin…
arxiv.org17 days agoView details
RouteSparse: Input-Conditional Pattern Routing for Budgeted Long-Context Prefilling
arXiv:2608.29058v1 Announce Type: new Abstract: Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and estimates that pattern's sparse indices for every prompt. This design is efficient, but it ass…
arxiv.org17 days agoView details
The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
arXiv:2608.28930v1 Announce Type: new Abstract: Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal over…
arxiv.org17 days agoView details
Models
View all models →time-series-forecasting · safetensors · time-series · forecasting
huggingface.co17 days ago832 ptsView details
OpenMOSS-Team/MOSS-Transcribe-Diarize
audio-text-to-text · transformers · safetensors · moss_transcribe_diarize
huggingface.co18 days ago406 ptsView details
sentence-similarity · PyLate · safetensors · modernbert
huggingface.co18 days ago44 ptsView details
text-generation · litert-lm · litert · litertlm
huggingface.co18 days ago3 ptsView details
VaultLevel6/beaverai_gguf_backups
gguf · endpoints_compatible · region:us
huggingface.co18 days agoView details
mradermacher/Puro-2B-Base-i1-GGUF
transformers · gguf · base-model
huggingface.co18 days agoView details
ishikaa/acquisition_student_AS_format_numina_qwen7b
text-generation · transformers · safetensors · qwen2
huggingface.co18 days agoView details
text-generation · transformers · safetensors · gguf
huggingface.co18 days agoView details
VertexAGI/prism-roleplay-1.5-small
text-generation · mlx · safetensors · gguf
huggingface.co18 days agoView details
VertexAGI/prism-roleplay-1-small
text-generation · mlx · safetensors · gguf
huggingface.co18 days agoView details
Open source
View all open source →<details open> metal : add fa-vec tunings for M1 Ultra (#28088) * metal : add fa-vec tunings for M1 Ultra * metal : move M1 Ultra tunings after M1 Max section * metal : remove duplicate blank line </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.…
github.com17 days agoView details
modelcontextprotocol/servers 2026.8.31
# Release : v2026.8.31 ## Updated packages - @modelcontextprotocol/server-filesystem@2026.8.31 - @modelcontextprotocol/server-memory@2026.8.31 - @modelcontextprotocol/server-sequential-thinking@2026.8.31 - @modelcontextprotocol/server-everything@2026.8.31
github.com17 days agoView details
<details open> vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (#28032) * vulkan: add top-k radix sort shader for k >= 1024 * add Qwen 3.8 Flash Next top-k tests * add top-k qsa fusion * clean up code </details> **Website:** - <https://llama.app> **Attestations:…
github.com18 days agoView details