Skip to content

Archive / 2026-08-11

August 11, 2026

  1. llama.cpp

    llama.app1 month ago364 ptsView detailsJoin discussion

  2. Grok Bot

    x.ai1 month ago338 ptsView detailsJoin discussion

  3. Lean Eval for Alignment on Faithfulness

    millenniumresearch.ai1 month ago103 ptsView detailsJoin discussion

  4. Why your Amazon order confirmation emails have become so unhelpful

    Earlier this summer, Amazon customers began noticing that emails related to their online orders looked sparse: Order confirmation emails didn't name specific items anymore, and instead listed only item categories. "Your Beauty item is confirmed!" an email about my retainer cleaning tablets read. Shoppers have posted o…

    theverge.com1 month ago78 ptsView detailsJoin discussion

  5. The whole of PyTorch on one page

    tensor.khalilli.ai1 month ago35 ptsView detailsJoin discussion

  6. AI Is Solving CTF Challenges in Minutes

    simulationslabs.com1 month ago21 ptsView detailsJoin discussion

  7. Show HN: Tokyo Trains

    greggman.github.io1 month ago12 ptsView detailsJoin discussion

  8. Never Bet Against Elon

    blog.oxplot.com1 month ago11 ptsView detailsJoin discussion

  9. Primary source

    How RingCentral builds AI-native work from engineering to ops

    See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.

    openai.com1 month agoView details

  10. Primary source

    Testing ads in ChatGPT

    OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.

    openai.com1 month agoView details

  11. Primary source

    Daybreak models are now available on AWS

    OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

    openai.com1 month agoView details

  12. Primary source

    Advancing AMIE towards expert-level audio-visual clinical consultations

    Health & Bioscience

    research.google1 month agoView details

  13. Primary source

    Thinking of ACE? We Can Do It with Fewer Tokens

    huggingface.co1 month agoView details

  14. The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model

    LTX-2.5 brings frontier video generation to local NVIDIA hardware: 6.8-second clips, native multishot, day-one ComfyUI, open weights. The post The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model appeared first on MarkTechPost.

    marktechpost.com1 month agoView details

  15. Building and Validating a Quantitative Trading Strategy with OctoBot, Walk-Forward Backtesting, Parameter Optimization, and Interactive Analysis

    In this tutorial, we build a complete quantitative backtesting workflow with OctoBot and OctoBot-Script while keeping the environment isolated from Colab’s preinstalled dependencies. We configure a rule-based trading strategy that combines RSI-based oversold signals, EMA trend confirmation, and ATR-driven adaptive sto…

    marktechpost.com1 month agoView details

  16. webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware

    webAI has released TwIL-LM, a family of formal-logic models at 1.7B and 3B parameters that translate English into first-order logic and check whether conclusions follow from premises. The 3B runs on CPU or 4GB of VRAM; the 1.7B downloads at 1.06GB. Both ship under a non-commercial license. The model card also shows th…

    marktechpost.com1 month agoView details

  17. Implementing a MiniMax-H3 Multimodal Video and Audio Generation Pipeline with ComfyUI APIs

    In this comprehensive guide, we demonstrate how to implement a complete, programmable MiniMax-H3 multimodal generation pipeline. By leveraging ComfyUI as a headless backend, we walk through setting up an automated inference environment that handles hardware profiling, model weight downloading, dynamic graph constructi…

    marktechpost.com1 month agoView details

  18. Saber denies replacing Rideshare Stimulator’s writers with ChatGPT

    After a former lead writer claimed Saber "replaced me with ChatGPT," CEO Matthew Karch now claims, "Neither Saber nor Unigine have replaced any writers with AI," for the Rideshare "Stimulator" game announced last month, developed by Unigine. The writer, Stella Sacco, says differently, however, posting on Bluesky that…

    theverge.com1 month agoView details

  19. ChatGPT and Gemini both just passed 1 billion users

    That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milest…

    theverge.com1 month agoView details

  20. Another OpenAI executive takes off

    Brad Lightcap, OpenAI's special projects lead and the company's former COO, announced his departure after an eight-year stint at the AI lab. In an internal memo he later posted to X, Lightcap told colleagues he'd be starting "something new." "Over the last few months, I've been focused on the next horizon and what wou…

    theverge.com1 month agoView details

  21. Made by Google 2026: all the Pixel news and announcements

    On August 12, 2026, Google revealed a bunch of new Pixel devices. The colorful Pixel 11 lineup comes with upgraded cameras and performance, with the Pro models offering a built-in LED ring that lights up for Google’s Gemini AI and other features. Google also showed off its next-gen Pixel Fold featuring thinner bezels,…

    theverge.com1 month agoView details

  22. Apple could help you prove your iPhone photos aren’t deepfakes

    Apple is seemingly developing an iOS feature that can verify when a photograph was taken using an iPhone camera. 9to5Mac reports that the iOS 27 beta 5 includes code references for an "Apple Reference Image" system that can embed provenance metadata into iPhone photographs at the point of capture - enabling users to p…

    theverge.com1 month agoView details

  23. ‘Zoomsday’ hack uncovered using fewer than 20 AI prompts

    Zoom has patched a major security vulnerability that could allow an attacker to hijack anyone's device during a meeting. In a blog post on Tuesday, researchers at A Security say they uncovered the flaw using "fewer than 20 prompts on publicly available AI models," as reported earlier by Wired. The exploit involved Zoo…

    theverge.com1 month agoView details

  24. Spotify says it won’t recommend music from ‘AI Personas’

    Spotify will soon label AI artists and remove their music from your recommendations. The change, which will start rolling out in mid-September, means you'll see an "AI Persona" badge on an artist's profile across the app if they do "not represent a real person." The music streaming platform will allow artists to discl…

    theverge.com1 month agoView details

  25. A Zoom Screen-Sharing Bug Let Anyone Take Over Other Devices on a Call

    Researchers say it took fewer than 20 prompts for a public AI tool to find a flaw (now fixed) allowing anyone on a Zoom call to hijack another participants’ device.

    wired.com1 month agoView details

  26. Claude will apply invisible watermarks to AI text and images

    Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a…

    theverge.com1 month agoView details

  27. The AI takeover of mathematics has begun

    Mathematician James Maynard has spent a lot of time this past year "soul searching." A professor at the University of Oxford and winner of the prestigious Fields Medal, Maynard told The Verge he's been grappling with the future of his field as the traditionally slow-moving discipline hurries to adapt to AI. Days befor…

    theverge.com1 month agoView details

  28. A New Trick Reveals AI Models’ Inner Thoughts

    Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.

    wired.com1 month agoView details

  29. AI Is Dead. Organoids Are Alive

    Mini human brains are being grown in labs all over the world. Soon, they could outthink neural networks.

    wired.com1 month agoView details

  30. AI Is Helping Solve the Intricate Genetic Puzzle of Schizophrenia

    Recent findings provide one of the most detailed pictures to date of the genetic architecture of schizophrenia, opening up new avenues for research into the disorder.

    wired.com1 month agoView details

  31. TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

    arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We intr…

    arxiv.org1 month agoView details

  32. SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

    arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat RAG diaries and Markdown logs optimize needle retrieval but under-serve currency, provenance, and ordered temporal rea…

    arxiv.org1 month agoView details

  33. Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

    arXiv:2608.10315v1 Announce Type: new Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a…

    arxiv.org1 month agoView details

  34. Towards an Argumentative Foundation for Evaluative AI

    arXiv:2608.07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) arg…

    arxiv.org1 month agoView details

  35. The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era

    arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a curriculum-framework paper, grounded in a structured narrative review of labor-market and software-engineering evidence and illustrate…

    arxiv.org1 month agoView details

  36. A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

    arXiv:2608.10939v1 Announce Type: new Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference po…

    arxiv.org1 month agoView details

  37. TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

    arXiv:2608.11044v1 Announce Type: new Abstract: Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengt…

    arxiv.org1 month agoView details

  38. Towards Researcher Agents for Knowledge-Graph Question Answering

    arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and seman…

    arxiv.org1 month agoView details

  39. Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

    arXiv:2608.07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking rich comparative s…

    arxiv.org1 month agoView details

  40. Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs

    arXiv:2608.07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executions are non-deterministic, and AI workloads bring high-dimensional inputs and non-linear operations that defeat fuzzing and symbolic…

    arxiv.org1 month agoView details

  41. Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility

    arXiv:2608.07476v1 Announce Type: new Abstract: We develop a formal framework for constructing canonical interpretations from plural structure theories. A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose admissible interpretation family collects all gl…

    arxiv.org1 month agoView details

  42. Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

    arXiv:2608.10678v1 Announce Type: new Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3)…

    arxiv.org1 month agoView details

  43. Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

    arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion…

    arxiv.org1 month agoView details

  44. Legal Responsibilities Using Autonomous Agents For Artificial Intelligence

    arXiv:2608.08022v1 Announce Type: new Abstract: Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant criminal or negligent d…

    arxiv.org1 month agoView details

  45. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

    arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The worl…

    arxiv.org1 month agoView details

  46. Emotion in an active inference model of human driving

    arXiv:2608.07480v1 Announce Type: new Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It has been successfully applied across biological and artificial systems, including recent work on human driving. However,…

    arxiv.org1 month agoView details

  47. Training Variable Long Sequences with Data-Centric Parallel

    arXiv:2608.07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches use static configurations that cause workload imbalance low efficiency,…

    arxiv.org1 month agoView details

  48. An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

    arXiv:2608.07542v1 Announce Type: new Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of whichever metric they optimise rather than testing the hypotheses that motivate the experiments. We address this structur…

    arxiv.org1 month agoView details

  49. CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

    arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasonin…

    arxiv.org1 month agoView details

  50. VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

    arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed char…

    arxiv.org1 month agoView details

  51. How Robust Are LLMs to Vietnamese Dialects?

    arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through d…

    arxiv.org1 month agoView details

  52. Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

    arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show thi…

    arxiv.org1 month agoView details

  53. Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

    arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or position-dependent rotations into Transformer representa…

    arxiv.org1 month agoView details

  54. Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

    arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape tru…

    arxiv.org1 month agoView details

  55. Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems

    arXiv:2608.10216v1 Announce Type: new Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question: "Does this text still mean the same thin…

    arxiv.org1 month agoView details

  56. Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

    arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $\beta$-fraction of shifted target traffic with at most an $\alpha$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Fl…

    arxiv.org1 month agoView details

  57. ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

    arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventi…

    arxiv.org1 month agoView details

  58. Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

    arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities,…

    arxiv.org1 month agoView details

  59. Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

    arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. We show it is functional. A 12-layer model computes in two phases. Throu…

    arxiv.org1 month agoView details

  60. Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation

    arXiv:2608.07949v1 Announce Type: new Abstract: Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneous sources, but prov…

    arxiv.org1 month agoView details

  61. Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

    arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems should be built. We at…

    arxiv.org1 month agoView details

  62. Simplex Relaxation for Discrete Diffusion

    arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask whether its training objective and reverse tr…

    arxiv.org1 month agoView details

  63. AndroidReality: How Far Are Mobile Agents from the Real World?

    arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions. In this work, we introduce AndroidReal…

    arxiv.org1 month agoView details

  64. From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems

    arXiv:2608.07627v1 Announce Type: new Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge…

    arxiv.org1 month agoView details

  65. Most biomedical publications show signs of LLM-assisted writing

    arXiv:2608.10715v1 Announce Type: new Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decis…

    arxiv.org1 month agoView details

  66. Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems

    arXiv:2608.07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and error propagation in…

    arxiv.org1 month agoView details

  67. TRIBE: Predicting Team Performance via Communication Behavior Ensembles

    arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metr…

    arxiv.org1 month agoView details

  68. QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing

    arXiv:2608.07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible complexity scope. We pre…

    arxiv.org1 month agoView details

  69. Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes

    arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply. The shape recurs: an origin's request-rate cap split across its edge locations, a licensed throughput cap ac…

    arxiv.org1 month agoView details

  70. MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales

    arXiv:2608.10974v1 Announce Type: new Abstract: Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often along with reasoning on why those methods were chosen. We introduce MUSE (Mining Underlying Scientific Explanations),…

    arxiv.org1 month agoView details

  71. Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

    arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic systems, a multi-component form of self…

    arxiv.org1 month agoView details

  72. The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

    arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes…

    arxiv.org1 month agoView details

  73. Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns

    arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by fixed rules. Here, we present Agent-MD,…

    arxiv.org1 month agoView details

  74. Contextual Value Alignment via Multilayer Combinatorial Fusion

    arXiv:2608.07642v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified re…

    arxiv.org1 month agoView details

  75. IntelliAudit: Using Large Language Models to Evaluate Audit Controls

    arXiv:2608.07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational…

    arxiv.org1 month agoView details

  76. REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

    arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a budget of at most 32B parameters and no model fine-tuning. Our system combines structured chain-of-thought reasoning, rela…

    arxiv.org1 month agoView details

  77. From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

    arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: e…

    arxiv.org1 month agoView details

  78. Protecting patient privacy in clinical foundation models: Technical and legal perspectives

    arXiv:2608.07705v1 Announce Type: new Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As deployment expands, privacy risk increasingly arises from model-mediated leakage, yet its prevalence and severity remain poorly qua…

    arxiv.org1 month agoView details

  79. GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

    arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for…

    arxiv.org1 month agoView details

  80. Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus

    arXiv:2608.10688v1 Announce Type: new Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking h…

    arxiv.org1 month agoView details

  81. Data Attribution of Emergent Misalignment with Persona Features

    arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligne…

    arxiv.org1 month agoView details

  82. Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

    arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally…

    arxiv.org1 month agoView details

  83. Mitigating Context Interference for Reliable and Efficient Search Agents

    arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts of multi-turn search agents are lengthy and complex. For example, the retrieved set of doc…

    arxiv.org1 month agoView details

  84. An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

    arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma…

    arxiv.org1 month agoView details

  85. MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

    arXiv:2608.10459v1 Announce Type: new Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models. Input-only encoder detectors are suitable…

    arxiv.org1 month agoView details

  86. When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

    arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean a…

    arxiv.org1 month agoView details

  87. Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

    arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a member of the pre-training corpus of a…

    arxiv.org1 month agoView details

  88. EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

    arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper prese…

    arxiv.org1 month agoView details

  89. What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

    arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the…

    arxiv.org1 month agoView details

  90. ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

    arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers: responses may be partly correct, incomplete, or…

    arxiv.org1 month agoView details

  91. SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution

    arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may expose private user information and incu…

    arxiv.org1 month agoView details

  92. ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration

    arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documentation inspection,…

    arxiv.org1 month agoView details

  93. SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control

    arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a stable physical object but a latent and temporally evolving attention state. In this work, we…

    arxiv.org1 month agoView details

  94. The Illusion of Cross-Lingual Safety in Low-Resource Languages

    arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross…

    arxiv.org1 month agoView details

  95. JustLLMGRPO: Radiographic Control for Chest X-Ray Generation

    arXiv:2608.08046v1 Announce Type: new Abstract: Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain adaptation. We sho…

    arxiv.org1 month agoView details

  96. From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

    arXiv:2608.11171v1 Announce Type: new Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mecha…

    arxiv.org1 month agoView details

  97. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

    arXiv:2608.11200v1 Announce Type: new Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while b…

    arxiv.org1 month agoView details

  98. When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

    arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four…

    arxiv.org1 month agoView details

  99. Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

    arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. However, existing benchmarks provide limited assessment of whether LLMs can faithfully perform multi-hop reasoning chains a…

    arxiv.org1 month agoView details

  100. Back to the Future: A workbook time machine for spread sheet creation benchmarks

    arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public workbook corpora,…

    arxiv.org1 month agoView details

  101. Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning

    arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generation. Tool augmentat…

    arxiv.org1 month agoView details

  102. SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

    arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent sys…

    arxiv.org1 month agoView details

  103. VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge

    arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse r…

    arxiv.org1 month agoView details

  104. Thought-Level Beam Search for Reasoning

    arXiv:2608.08020v2 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time…

    arxiv.org1 month agoView details

  105. Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities

    arXiv:2608.08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nev…

    arxiv.org1 month agoView details

  106. LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

    arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user re…

    arxiv.org1 month agoView details

  107. PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

    arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively under…

    arxiv.org1 month agoView details

  108. The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

    arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid masking distorts the model's underlying probability distribution, often biasing generation toward vali…

    arxiv.org1 month agoView details

  109. Multimodal Item Parameter Estimation using Simulated Response Probabilitie

    arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a l…

    arxiv.org1 month agoView details

  110. Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

    arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies; in contrast, we estimate corpus ratios…

    arxiv.org1 month agoView details

  111. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

    arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capa…

    arxiv.org1 month agoView details

  112. On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

    arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we in…

    arxiv.org1 month agoView details

  113. Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

    arXiv:2608.11008v1 Announce Type: new Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbaggin…

    arxiv.org1 month agoView details

  114. myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

    arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by nat…

    arxiv.org1 month agoView details

  115. Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

    arXiv:2608.10627v1 Announce Type: new Abstract: Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims before checking each one. Decomposition itself is treated as a neutral preprocessing step. We show it is not: a decompos…

    arxiv.org1 month agoView details

  116. Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

    arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how th…

    arxiv.org1 month agoView details

  117. Attention-Path Fragility as an Uncertainty Signal in Large Language Models

    arXiv:2608.11138v1 Announce Type: new Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwor…

    arxiv.org1 month agoView details

  118. Divergent Response Modes in Frontier Language Models Under Steering Pressure

    arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six…

    arxiv.org1 month agoView details

  119. TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

    arXiv:2608.07917v2 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accur…

    arxiv.org1 month agoView details

  120. Multiclass Sentiment Analysis for Identifying Political Viewpoints

    arXiv:2608.11049v1 Announce Type: new Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) t…

    arxiv.org1 month agoView details

  121. CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

    arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement with CausalNav, a controller built around a signed, action-conditioned transition graph over i…

    arxiv.org1 month agoView details

  122. CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

    arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structu…

    arxiv.org1 month agoView details

  123. The Authority Expectancy Effect in Multi-User Conflict

    arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemini, GPT, Grok) and t…

    arxiv.org1 month agoView details

  124. MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

    arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level task co…

    arxiv.org1 month agoView details

  125. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

    arXiv:2608.10916v1 Announce Type: new Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive human-annotated ground truth, or rely on LLM j…

    arxiv.org1 month agoView details

  126. The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

    arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bi…

    arxiv.org1 month agoView details

  127. CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity

    arXiv:2608.07965v1 Announce Type: new Abstract: Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecurity education. Generative AI agents offer a path to delivering such experiences adaptively at scale, but introduce well-documented ri…

    arxiv.org1 month agoView details

  128. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

    arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final e…

    arxiv.org1 month agoView details

  129. TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

    arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce TelemetrySuffBench, a controlled benchmark that separates failure detection, fault-origin localiza…

    arxiv.org1 month agoView details

  130. GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering

    arXiv:2608.07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical r…

    arxiv.org1 month agoView details

  131. Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

    arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions…

    arxiv.org1 month agoView details

  132. Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

    arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the offi…

    arxiv.org1 month agoView details

  133. X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

    arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state p…

    arxiv.org1 month agoView details

  134. Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

    arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work,…

    arxiv.org1 month agoView details

  135. NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

    arXiv:2608.07530v1 Announce Type: new Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. Howe…

    arxiv.org1 month agoView details

  136. KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs

    arXiv:2608.07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different questions. In this work, we…

    arxiv.org1 month agoView details

  137. When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

    arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private…

    arxiv.org1 month agoView details

  138. TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

    arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc ed…

    arxiv.org1 month agoView details

  139. Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

    arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. Understanding these relationships is important for model provenance…

    arxiv.org1 month agoView details

  140. TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

    arXiv:2608.07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an…

    arxiv.org1 month agoView details

  141. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

    arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free qua…

    arxiv.org1 month agoView details

  142. Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

    arXiv:2608.07474v1 Announce Type: new Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load.…

    arxiv.org1 month agoView details

  143. Assessing Reliability of BERT-Based Models on Question Answering Tasks

    arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language processing (NLP), particularly those based on…

    arxiv.org1 month agoView details

  144. REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

    arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps propagate to an in…

    arxiv.org1 month agoView details

  145. The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes

    arXiv:2608.07566v1 Announce Type: new Abstract: We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework encodes a scene into coefficients of a fixed symmetric matrix basis, assembles them into a Lie algebra element, and exponentiates the result to a Riemannian or Lor…

    arxiv.org1 month agoView details

  146. Controlled Memory Interference in Continual LLM Agents

    arXiv:2608.07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise, or interfere with…

    arxiv.org1 month agoView details

  147. Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

    arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reaso…

    arxiv.org1 month agoView details

  148. ReLTEx: Reliable LLM-based Taxonomy Expansion

    arXiv:2608.10970v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansions often leads to noi…

    arxiv.org1 month agoView details

  149. Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025

    arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language…

    arxiv.org1 month agoView details

  150. When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

    arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically lar…

    arxiv.org1 month agoView details

  1. MiniMaxAI/MiniMax-H3

    image-text-to-video · minimax-h3 · diffusers · safetensors

    huggingface.co1 month ago5400 ptsView details

  2. Lightricks/LTX-2.5

    image-to-video · diffusion-single-file · image-to-video · text-to-video

    huggingface.co1 month ago4153 ptsView details

  3. meta-models/Muse-Glimmer-30B

    image-text-to-text · transformers · safetensors · muse_glimmer

    huggingface.co1 month ago1791 ptsView details

  4. lightx2v/Minimax-h3-Turbo

    image-to-video · diffusers · t2v · i2v

    huggingface.co1 month ago891 ptsView details

  5. nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

    safetensors · en · arxiv:2410.17196

    huggingface.co1 month ago392 ptsView details

  6. TenStrip/10Eros-Max

    image-text-to-video · text-to-video · image-text-to-video · image-to-video

    huggingface.co1 month ago345 ptsView details

  7. SexGod1979/PinkCherry_MiniMax-H3

    text-to-video · transformers · minimax-h3 · text-to-video

    huggingface.co1 month ago337 ptsView details

  8. inclusionAI/Ling-3.0-tiny

    text-generation · safetensors · bailing_hybrid · text-generation

    huggingface.co1 month ago346 ptsView details

  9. nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

    text-generation · transformers · safetensors · nemotron_h

    huggingface.co1 month ago331 ptsView details

  10. drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

    text-to-video · minimax-h3 · lora · adapter

    huggingface.co1 month ago327 ptsView details

  11. nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

    text-generation · transformers · safetensors · nemotron_h

    huggingface.co1 month ago172 ptsView details

  12. Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot

    image-text-to-video · diffusers · text-to-video · image-to-video

    huggingface.co1 month ago176 ptsView details

  13. Cactus-Compute/needle2

    text-generation · cactus-needle · needle · tool-calling

    huggingface.co1 month ago161 ptsView details

  14. Motif-Technologies/Motif-3

    text-generation · transformers · safetensors · Motif

    huggingface.co1 month ago136 ptsView details

  15. SyzygyResearch/Mach-1-Additive-35B

    qwen3_5_moe · qwen · mach-1

    huggingface.co1 month ago123 ptsView details

  16. peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF

    image-text-to-text · gguf · llama.cpp · qwen3_6

    huggingface.co1 month ago64 ptsView details

  17. Sachin21112004/distilbart-news-summarizer

    summarization · transformers · pytorch · jax

    huggingface.co1 month ago21 ptsView details

  18. jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame-hf

    automatic-speech-recognition · transformers · safetensors · qwen3_asr

    huggingface.co1 month ago2 ptsView details

  19. sullivan1502/base-zone-grpo

    text-generation · transformers · safetensors · llama

    huggingface.co1 month agoView details

  20. mobilint/whisper-large-v3-turbo

    automatic-speech-recognition · transformers · safetensors · mobilint-whisper

    huggingface.co1 month agoView details

  21. naver-ellm/HyperCLOVAX-SEED-Text-Instruct-0.5B-GGUF

    gguf · base_model:naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-0.5B · base_model:quantized:naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-0.5B

    huggingface.co1 month agoView details

  22. mobilint/whisper-medium

    automatic-speech-recognition · transformers · safetensors · mobilint-whisper

    huggingface.co1 month agoView details

  23. mobilint/whisper-small

    automatic-speech-recognition · transformers · safetensors · mobilint-whisper

    huggingface.co1 month agoView details

  1. ggml-org/llama.cpp b10369

    <details open> mtmd: support pocket-tts (#26871) * adapt the api * text model ok * working impl, need verify and clean up * mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample was built as one…

    github.com1 month agoView details

  2. openai/openai-python v3.0.0

    ## [3.0.0](https://github.com/openai/openai-python/compare/v2.54.0...v3.0.0) (2026-08-12) ### ⚠ BREAKING CHANGES * **api:** HTTPX2 is now the default HTTP client, and `httpx` is no longer installed automatically. Applications using custom HTTPX clients, transports, or configurat…

    github.com1 month agoView details

  3. ggml-org/llama.cpp b10362

    <details open> tests : disable backend sampler hip multi output (#26878) * test-backend-sampler: skip multi_output_sampling_chain on HIP The new multi_output_sampling_chain test uses top_k, whose backend probs path needs CUB (unavailable on HIP), so sampled_probs is null and the…

    github.com1 month agoView details

  4. langchain-ai/langchain langchain-anthropic==1.5.5

    Changes since langchain-anthropic==1.5.4 release(anthropic): 1.5.5 (#39597) fix(anthropic): report reasoning tokens in usage metadata (#39590) fix(anthropic): fix `KeyError` on rename in Claude file-tool middleware (#39293) chore(model-profiles): refresh model profile data (#392…

    github.com1 month agoView details

  5. langchain-ai/langchain langchain==1.3.15

    Changes since langchain==1.3.14 release(langchain): 1.3.15 (#39595) feat(langchain): expose `trace_policy` on `AgentMiddleware` (#38910) chore(langchain): fix type errors in tests (#39589) chore: bump h2 from 4.3.0 to 4.4.1 in /libs/langchain_v1 (#39324) fix(langchain): preserve…

    github.com1 month agoView details

  6. openai/openai-python v2.54.0

    ## [2.54.0](https://github.com/openai/openai-python/compare/v2.53.0...v2.54.0) (2026-08-11) ### Features * **api:** Add new Responses model identifiers ([#3595](https://github.com/openai/openai-python/issues/3595)) ([0652787](https://github.com/openai/openai-python/commit/065278…

    github.com1 month agoView details

  7. langchain-ai/langchain langchain-core==1.5.4

    Changes since langchain-core==1.5.3 release(core): 1.5.4 (#39592) fix(core): compat with pydantic 2.14 (#39328) fix(core): stop StructuredPrompt from mutating caller kwargs (#39174) fix(core): preserve flat tool args schema for `RootModel` runnables (#39307) fix(core): close int…

    github.com1 month agoView details

  8. ggml-org/llama.cpp b10361

    <details open> model : fix SWA not being enabled for EXAONE 4.5 (#26848) * model : fix SWA not being enabled for EXAONE 4.5 load_arch_hparams tests `hparams.n_layer() == 64` before LLM_KV_NEXTN_PREDICT_LAYERS has been read. n_layer() returns n_layer_all - n_layer_nextn and n_lay…

    github.com1 month agoView details

  9. vllm-project/vllm v0.27.1

    This is a patch release on top of v0.27.0. - Support quantized DSpark Markov heads (#50424)

    github.com1 month agoView details