Skip to content

Archive / 2026-08-23

August 23, 2026

  1. Training AI to Paint with Code

    surya.website25 days ago206 ptsView detailsJoin discussion

  2. Erik Brynjolfsson says an AI "job apocalypse" is unlikely

    wpintelligence.washingtonpost.com25 days ago38 ptsView detailsJoin discussion

  3. Scientific Data Analysis with LabPlot in Python: Signal Processing, Spectral Peak Fitting, Visualization, and Batch Automation

    In this tutorial, we explore a LabPlot-inspired scientific data analysis workflow in Python while preserving the structure and terminology of LabPlot’s aspect tree, analysis kernels, plotting system, and project model. We build reusable components to import tabular data, compute descriptive statistics, smooth and diff…

    marktechpost.com25 days agoView details

  4. Harvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work

    Harvey's first post-trained model nearly doubles LAB task completion, but only one benchmark number survives independent verification today The post Harvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work appeared first on MarkTechPost.

    marktechpost.com25 days agoView details

  5. Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

    FreeToken splits MoE cache misses between PCIe fills and CPU execution using measured bandwidths, unlocking frontier models locally The post Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU appeared first on MarkTechPost.

    marktechpost.com25 days agoView details

  6. Building an End-to-End Document Intelligence Pipeline with deepDoctection

    Build an end-to-end document intelligence pipeline with deepDoctection. This tutorial covers configuring layout analysis, DocTR OCR, and table extraction, while demonstrating how to implement custom services for entity recognition and generate structured JSONL data for your RAG workflows. The post Building an End-to-E…

    marktechpost.com26 days agoView details

  7. Vercel Introduces ‘Is Agentic’, a Free Agent-Readiness Scoring Tool That Audits Public Websites Using Ora’s 100+ Checks

    Vercel and Ora launched Is Agentic, a free audit scoring website readiness for AI agents across 118 checks. The post Vercel Introduces ‘Is Agentic’, a Free Agent-Readiness Scoring Tool That Audits Public Websites Using Ora’s 100+ Checks appeared first on MarkTechPost.

    marktechpost.com26 days agoView details

  8. PromptResponse: Optimizing Prompts for LLM Coding Tasks

    arXiv:2608.21074v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unicode{x00BB}$, a controlled study examining…

    arxiv.org25 days agoView details

  9. When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems

    arXiv:2608.20371v1 Announce Type: new Abstract: A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this claim head-to-head and find that the honest answer is: it depends on the intent space. On full ATIS and CLINC150 we compare a fine-tune…

    arxiv.org25 days agoView details

  10. Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

    arXiv:2608.20361v1 Announce Type: new Abstract: Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three approaches fail in the same way: each treats…

    arxiv.org25 days agoView details

  11. TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

    arXiv:2608.20360v1 Announce Type: new Abstract: We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a product-only degree-3 branch that multiplies three…

    arxiv.org25 days agoView details

  12. Source-Free MT Evaluation Is Not MT Evaluation

    arXiv:2608.20925v1 Announce Type: new Abstract: Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practical norm, even though…

    arxiv.org25 days agoView details

  13. When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

    arXiv:2608.20400v1 Announce Type: new Abstract: Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurally indirect prerequi…

    arxiv.org25 days agoView details

  14. LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

    arXiv:2608.20530v1 Announce Type: new Abstract: Speculative decoding accelerates language-model inference by drafting future tokens that the target model verifies in parallel. A diffusion-style block head such as DFlash is an attractive drafter, predicting an entire block of future tokens in one forward pass. However,…

    arxiv.org25 days agoView details

  15. Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

    arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speec…

    arxiv.org25 days agoView details

  16. Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift

    arXiv:2608.21043v1 Announce Type: new Abstract: Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially relevant in social-engineering fraud detection, where attackers can preserve malicious intent w…

    arxiv.org25 days agoView details

  17. Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol

    arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observation…

    arxiv.org25 days agoView details

  18. Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context

    arXiv:2608.20807v1 Announce Type: new Abstract: Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literature-informed environmental priors can serve as an au…

    arxiv.org25 days agoView details

  19. Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness

    arXiv:2608.20389v1 Announce Type: new Abstract: A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small scale this selection happens in context: the LLM planner chooses among skill representations exposed in its system prompt, without an…

    arxiv.org25 days agoView details

  20. Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

    arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing f…

    arxiv.org25 days agoView details

  21. STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction

    arXiv:2608.20831v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought (CoT) models perform well on this task, dis…

    arxiv.org25 days agoView details

  22. ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding

    arXiv:2608.20405v1 Announce Type: new Abstract: Persuasive argument generation requires modeling audience beliefs, rhetorical strategies, and factual grounding. Despite recent advancements, existing methods remain largely audience-agnostic and fail to integrate strategy selection to improve persuasiveness. To bridge t…

    arxiv.org25 days agoView details

  23. The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

    arXiv:2608.20353v1 Announce Type: new Abstract: Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagnostic framework that d…

    arxiv.org25 days agoView details

  24. A Temporal Planning Approach for Intelligent Flood Response

    arXiv:2608.20510v1 Announce Type: new Abstract: Effective response to multiple, simultaneously flooded areas requires coordinating appropriate actions in the correct temporal order, under severe resource constraints. Automated planning provides a foundation for addressing this challenge by generating time-aware schedu…

    arxiv.org25 days agoView details

  25. Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

    arXiv:2608.20362v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an…

    arxiv.org25 days agoView details

  26. ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora

    arXiv:2608.20369v1 Announce Type: new Abstract: Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constructing a reporting te…

    arxiv.org25 days agoView details

  27. An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

    arXiv:2608.20373v1 Announce Type: new Abstract: Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiov…

    arxiv.org25 days agoView details

  28. VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

    arXiv:2608.20374v1 Announce Type: new Abstract: How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the desired affect as a continuous point (v*, a…

    arxiv.org25 days agoView details

  29. Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

    arXiv:2608.20755v1 Announce Type: new Abstract: Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and i…

    arxiv.org25 days agoView details

  30. World models of environment, agent and joint agent-environment systems

    arXiv:2608.20401v1 Announce Type: new Abstract: World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, states, latent or information states. We argue that there is a prior distinction: which channel…

    arxiv.org25 days agoView details

  31. Categorical AI phenomenology: A first-person approach

    arXiv:2608.20420v1 Announce Type: new Abstract: This paper develops a phenomenology-first approach to artificial consciousness by reframing consciousness as the subjective experience enacted through an agent's interface with the world. We shift the methodological focus to first-person structures, modeled mathematicall…

    arxiv.org25 days agoView details

  32. GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring

    arXiv:2608.20375v1 Announce Type: new Abstract: Tree-based speculative decoding raises the mean accepted tokens of standard speculative decoding by verifying multiple draft paths, and existing tree builders typically construct these paths through parent-conditioned expansion, where each child token is generated condit…

    arxiv.org25 days agoView details

  33. Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

    arXiv:2608.20385v1 Announce Type: new Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treat…

    arxiv.org25 days agoView details

  34. Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

    arXiv:2608.20387v1 Announce Type: new Abstract: While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open-ended instructions using in-the-wild aud…

    arxiv.org25 days agoView details

  35. SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL

    arXiv:2608.20630v1 Announce Type: new Abstract: SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering, ranking, retrieval, joining, and summarization. Despite their diverse APIs, these functions play only three relational roles: transforming individual rows, aggregating g…

    arxiv.org25 days agoView details

  36. Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

    arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to outdated, irrelevant, or misleading in…

    arxiv.org25 days agoView details

  37. Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

    arXiv:2608.20964v1 Announce Type: new Abstract: In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propos…

    arxiv.org25 days agoView details

  38. Scaling Unsupervised Word Alignment to Documents via Structural Constraints

    arXiv:2608.21023v1 Announce Type: new Abstract: Word alignment has traditionally been studied between sentences, but many cross-lingual tasks increasingly require correspondences across full documents. While recent multilingual embedding models can encode long inputs, we show that applying algorithms designed for sent…

    arxiv.org25 days agoView details

  39. RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

    arXiv:2608.20845v1 Announce Type: new Abstract: Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re-derives the meaning of raw corpus text and then throws that work away. Cheaper models do not close the gap: per-token prices have f…

    arxiv.org25 days agoView details

  40. Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts

    arXiv:2608.20490v1 Announce Type: new Abstract: AI ethics frameworks treat values such as fairness, transparency, and accountability as universal and uniformly operationalizable across contexts. We examined how 14 experts across 10 countries made sense of AI in practice, reinterpreted core values, and envisioned gover…

    arxiv.org25 days agoView details

  41. Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles

    arXiv:2608.20384v1 Announce Type: new Abstract: Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive accuracy and produce human-understandable explanations of the cues driving their decisions -- a dual objective that current high-ca…

    arxiv.org25 days agoView details

  42. Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation

    arXiv:2608.20794v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic forgetting, yet open-ended factual failures do not necessarily imply that the underlying facts have been erased. In this work, we iden…

    arxiv.org25 days agoView details

  43. Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

    arXiv:2608.20347v1 Announce Type: new Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are…

    arxiv.org25 days agoView details

  44. Self-Speculation for Faster Reasoning Models

    arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor fit for latency-sensitive and interacti…

    arxiv.org25 days agoView details

  45. Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing

    arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use th…

    arxiv.org25 days agoView details

  46. Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

    arXiv:2608.20349v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger disproportionate performance fluctuations. Moving beyond black-box optimization and coarse-grained templates, we present the first larg…

    arxiv.org25 days agoView details

  47. How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

    arXiv:2608.20350v1 Announce Type: new Abstract: Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high latency. We propose O…

    arxiv.org25 days agoView details

  48. DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

    arXiv:2608.20664v1 Announce Type: new Abstract: DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregis…

    arxiv.org25 days agoView details

  49. Decoupled Vision-Language System for Multimodal Understanding and Generation

    arXiv:2608.20382v1 Announce Type: new Abstract: We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one vision system and one language system, connected by cross-modal bridges. This design decou…

    arxiv.org25 days agoView details

  50. StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

    arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain kn…

    arxiv.org25 days agoView details

  51. Hadith computational science in the age of large language models: a critical narrative review

    arXiv:2608.20364v1 Announce Type: new Abstract: We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are met…

    arxiv.org25 days agoView details

  52. Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

    arXiv:2608.20365v1 Announce Type: new Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morp…

    arxiv.org25 days agoView details

  53. EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators

    arXiv:2608.20381v1 Announce Type: new Abstract: Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or…

    arxiv.org25 days agoView details

  54. Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI

    arXiv:2608.20393v1 Announce Type: new Abstract: Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables fine-tuning-free style control…

    arxiv.org25 days agoView details

  55. Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

    arXiv:2608.20396v1 Announce Type: new Abstract: Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting scalability across languages and populations. Her…

    arxiv.org25 days agoView details

  56. Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

    arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim on eight open-weig…

    arxiv.org25 days agoView details

  57. FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

    arXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task…

    arxiv.org25 days agoView details

  58. JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

    arXiv:2608.20607v1 Announce Type: new Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independen…

    arxiv.org25 days agoView details

  59. When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

    arXiv:2608.20627v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the trajectory. This paper introduces…

    arxiv.org25 days agoView details

  60. Sparse Token Routing in Efficient Transformers

    arXiv:2608.20632v1 Announce Type: new Abstract: Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that routes tokens through either lightweight o…

    arxiv.org25 days agoView details

  61. AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

    arXiv:2608.20634v1 Announce Type: new Abstract: Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realis…

    arxiv.org25 days agoView details

  62. MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees

    arXiv:2608.20636v1 Announce Type: new Abstract: Many text classification decisions are viable based on constituent excerpts alone. Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting such excerpts. We show that our approac…

    arxiv.org25 days agoView details

  63. Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails

    arXiv:2608.20647v1 Announce Type: new Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (strictly a function of tokens $i..n$) beats either alone and beats a fused self-attention representation for dependency…

    arxiv.org25 days agoView details

  64. Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation

    arXiv:2608.20804v1 Announce Type: new Abstract: Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on subject histories may insuff…

    arxiv.org25 days agoView details

  65. SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields

    arXiv:2608.20839v1 Announce Type: new Abstract: Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d. perturbations, which can be poorly align…

    arxiv.org25 days agoView details

  66. Ontology-Driven Structural Regularization for Document-Level Relation Extraction

    arXiv:2608.20856v1 Announce Type: new Abstract: Document-Level Relation Extraction (DocRE) relies heavily on costly manually annotated datasets, while large distant supervision resources such as DocRED distant remain underexploited due to noise. We show that a critical yet overlooked source of noise lies in structural…

    arxiv.org25 days agoView details

  67. KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

    arXiv:2608.20887v1 Announce Type: new Abstract: Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods ty…

    arxiv.org25 days agoView details

  68. ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction

    arXiv:2608.20920v1 Announce Type: new Abstract: Open-web future event prediction requires agents to distill reliable signals from noisy, redundant, and incomplete evidence. Existing retrieval/memory mechanisms directly feed retrieved information to agents or rely on simple memory functions such as storing and reusing…

    arxiv.org25 days agoView details

  69. DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

    arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steering prompts, the resulting answer-confide…

    arxiv.org25 days agoView details

  70. Continuous-Time Quantum Walks based Graph Neural Network

    arXiv:2608.20738v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Second, stacking layers…

    arxiv.org25 days agoView details

  71. LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine

    arXiv:2608.20402v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for linking Traditional Chinese Medicine (TCM…

    arxiv.org25 days agoView details

  72. Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

    arXiv:2608.20743v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel genera…

    arxiv.org25 days agoView details

  73. MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

    arXiv:2608.20927v1 Announce Type: new Abstract: Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. O…

    arxiv.org25 days agoView details

  74. SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers

    arXiv:2608.20802v1 Announce Type: new Abstract: Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty estimates that are structured, calibrated, and efficient. Bayesian and ensemble-based uncertainty estimates often require repeated stochastic inference [15, 26], whi…

    arxiv.org25 days agoView details

  75. Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

    arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially…

    arxiv.org25 days agoView details

  76. Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

    arXiv:2608.20825v1 Announce Type: new Abstract: Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. The safeguards that…

    arxiv.org25 days agoView details

  77. Foundation Models for Partial Causal Identification

    arXiv:2608.20841v1 Announce Type: new Abstract: This paper investigates the development of causal foundation models for bounding the effect of interventions and counterfactuals from observational data. We show that a canonical prior can be defined with full support over the space of structural causal models with discr…

    arxiv.org25 days agoView details

  78. TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding

    arXiv:2608.20844v1 Announce Type: new Abstract: Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as titles and images or missing from the catalog alto…

    arxiv.org25 days agoView details

  79. Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

    arXiv:2608.21019v1 Announce Type: new Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertai…

    arxiv.org25 days agoView details

  80. Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

    arXiv:2608.21021v1 Announce Type: new Abstract: Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether lightweight, edge-depl…

    arxiv.org25 days agoView details

  81. MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

    arXiv:2608.20853v1 Announce Type: new Abstract: Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this g…

    arxiv.org25 days agoView details

  82. Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

    arXiv:2608.20864v1 Announce Type: new Abstract: Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, th…

    arxiv.org25 days agoView details

  83. ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries

    arXiv:2608.20869v1 Announce Type: new Abstract: Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. Recent work on TS prediction have focused on flow matching supervised on straight linear paths that do not align with actual reaction trajectories. We p…

    arxiv.org25 days agoView details

  84. The Logic of Machine Self-Preservation

    arXiv:2608.20940v1 Announce Type: new Abstract: There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon known as instrument…

    arxiv.org25 days agoView details

  85. AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification

    arXiv:2608.20711v1 Announce Type: new Abstract: High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or t…

    arxiv.org25 days agoView details

  86. PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure

    arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's terminal-based coding age…

    arxiv.org25 days agoView details

  87. Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

    arXiv:2608.20378v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates th…

    arxiv.org25 days agoView details

  88. Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance

    arXiv:2608.20661v1 Announce Type: new Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to authoritative sources and auditable after the fac…

    arxiv.org25 days agoView details

  89. Why2Speak: Faithful Reasoning for Abstaining Action Policies

    arXiv:2608.20670v1 Announce Type: new Abstract: Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanation is useful only if it reflects the computation that produced the action. We study this problem through intervention timing in multi…

    arxiv.org25 days agoView details

  90. CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

    arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidat…

    arxiv.org25 days agoView details

  91. VortexChat: An agentic framework for autonomous multi-objective integrated photonic design

    arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inverse design offers an alternative, it remains constrained by expert supervision and a lack of end-to…

    arxiv.org25 days agoView details

  92. Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning

    arXiv:2608.20564v1 Announce Type: new Abstract: Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communication, particularly in hidden-profile settings where each agent holds only part of the evidence required for a correct decision. Existin…

    arxiv.org25 days agoView details

  93. Volumetric Radiology AI in the Era of Multimodal Large Language Models

    arXiv:2608.20549v1 Announce Type: new Abstract: Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch…

    arxiv.org25 days agoView details

  94. ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models

    arXiv:2608.20355v1 Announce Type: new Abstract: Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation,…

    arxiv.org25 days agoView details

  95. STCO: Conditional Neural Operators for Time-Dependent PDEs

    arXiv:2608.20477v1 Announce Type: new Abstract: Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs), but their future-state predictions are often conditioned only on observed states and static problem descriptors. For control or op…

    arxiv.org25 days agoView details

  96. Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

    arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choose among a growing menu of proposed interventions, but typically lack a principled basis for comparing them. Prior work tends to eva…

    arxiv.org25 days agoView details

  97. Intent Engine: Natural-Language Intent Translation for Intent-Driven Orchestration in the Compute Continuum

    arXiv:2608.20388v1 Announce Type: new Abstract: Microservice placement in the compute continuum is driven by low-level Service-level Objectives (SLOs), but requiring users to specify metric-level constraints creates an adoption barrier and increases misconfiguration risk. Although large language models (LLMs) can inte…

    arxiv.org25 days agoView details

  98. Dual-Cache Latent Space Communication between Heterogeneous Language Models

    arXiv:2608.20617v1 Announce Type: new Abstract: Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. They usually communicate by exchanging text, which puts autoregressi…

    arxiv.org25 days agoView details

  99. Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring

    arXiv:2608.20786v1 Announce Type: new Abstract: Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-response system, running an open-weights model under sovereignty constraints, and evaluate it against human-written bids the same org…

    arxiv.org25 days agoView details

  100. PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

    arXiv:2608.20757v1 Announce Type: new Abstract: We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual document-summary pairs, passage-based que…

    arxiv.org25 days agoView details

  101. A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

    arXiv:2608.20379v1 Announce Type: new Abstract: Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent…

    arxiv.org25 days agoView details

  102. Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

    arXiv:2608.20344v1 Announce Type: new Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summarie…

    arxiv.org25 days agoView details

  103. FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning

    arXiv:2608.20518v1 Announce Type: new Abstract: In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions, namely the communication topology, pe…

    arxiv.org25 days agoView details

  104. Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit

    arXiv:2608.20351v1 Announce Type: new Abstract: We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) o…

    arxiv.org25 days agoView details

  105. Environmental Slow AI: Design Principles for Generative Systems

    arXiv:2608.20398v1 Announce Type: new Abstract: Generative AI (genAI) systems produce cultural artefacts at scale, but they also reflect embedded cultural values through their design. Once identified, these values become open to deliberate reshaping. This position paper examines the maximalist values of current genera…

    arxiv.org25 days agoView details

  106. TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection

    arXiv:2608.20376v1 Announce Type: new Abstract: LLM agents can now generate realistic shilling profiles, fluent reviews, and coherent ratings at scale, systematically defeating recommender-system defenses. Text-only detectors that flag semantic drift in review embeddings are blind to graph structure and temporal coord…

    arxiv.org25 days agoView details

  107. Research Paper Quality Recognition Through Textual Feature Analysis

    arXiv:2608.20368v1 Announce Type: new Abstract: Knowledge and innovations are shaped by using the quality and credibility of the scientific research. Yet, distinguishing between impactful, high-quality work and flawed studies remains a challenge. This paper introduces a benchmark for classifying research papers into t…

    arxiv.org25 days agoView details

  108. Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

    arXiv:2608.20392v1 Announce Type: new Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation-as-Search (EaS), a f…

    arxiv.org25 days agoView details

  109. No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

    arXiv:2608.20938v1 Announce Type: new Abstract: Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment changes stem from va…

    arxiv.org25 days agoView details

  110. Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

    arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer…

    arxiv.org25 days agoView details

  111. Terminal Agents: A Survey of AI Agents in Command-Line Environments

    arXiv:2608.20485v1 Announce Type: new Abstract: Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bearing action--observa…

    arxiv.org25 days agoView details

  112. Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control

    arXiv:2608.20936v1 Announce Type: new Abstract: World models for continuous control are commonly trained for a fixed physical system and can degrade when known morphology parameters such as link lengths, masses, damping, and actuation change. Existing approaches often provide these parameters as conditioning informati…

    arxiv.org25 days agoView details

  113. Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

    arXiv:2608.20390v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We present Ansari, a deployed,…

    arxiv.org25 days agoView details

  114. ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research

    arXiv:2608.20391v1 Announce Type: new Abstract: Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce ImmigrationReason, a large-scale structured dataset d…

    arxiv.org25 days agoView details

  115. CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting

    arXiv:2608.20771v1 Announce Type: new Abstract: Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes critical evidence loss or noise inclusion, while over-confidence induced by progressive RL leads to hallucinated answers and redundant se…

    arxiv.org25 days agoView details

  116. Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

    arXiv:2608.20777v1 Announce Type: new Abstract: As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skepti…

    arxiv.org25 days agoView details

  117. SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

    arXiv:2608.20341v1 Announce Type: new Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substantial Functional Requ…

    arxiv.org25 days agoView details

  118. UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

    arXiv:2608.20918v1 Announce Type: new Abstract: Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work evaluates isolated mod…

    arxiv.org25 days agoView details

  119. Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

    arXiv:2608.20611v1 Announce Type: new Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the frozen SFT checkpo…

    arxiv.org25 days agoView details

  120. Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work

    arXiv:2608.20622v1 Announce Type: new Abstract: Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their own domain now costs an afternoon. The cost of reviewing and maintaining that code hasn't collapsed. Each solution drifts from the next; understanding one means rea…

    arxiv.org25 days agoView details

  121. When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

    arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general c…

    arxiv.org25 days agoView details

  122. Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation

    arXiv:2608.20797v1 Announce Type: new Abstract: Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at once, leading to substantial context over…

    arxiv.org25 days agoView details

  123. ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

    arXiv:2608.20735v1 Announce Type: new Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-scale teacher or ex…

    arxiv.org25 days agoView details

  124. Who Delegates to AI? Evidence from 53,000 Agent Configurations

    arXiv:2608.20425v1 Announce Type: new Abstract: A growing literature measures how far occupations are exposed to AI, but these measures capture where AI could perform tasks, not whether workers have adopted it. We propose a new layer of exposure, delegated exposure, which records whether a worker has committed a task…

    arxiv.org25 days agoView details

  125. Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

    arXiv:2608.20768v1 Announce Type: new Abstract: Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexamined. We propose a paire…

    arxiv.org25 days agoView details

  126. Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

    arXiv:2608.20953v1 Announce Type: new Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to requ…

    arxiv.org25 days agoView details

  127. Dynamic Context Scheduling: Learning Beyond the Static Universe

    arXiv:2608.20799v1 Announce Type: new Abstract: We study dynamic context scheduling as a training instrument for contextual re- inforcement learning. Rather than treating intra-episode context variation as a deployment reality, we treat it as a controlled shaping mechanism. Thereby, context evolves within each trainin…

    arxiv.org25 days agoView details

  1. AntResearch/4DAnyone

    video-to-video · safetensors · video-generation · multiview-video-generation

    huggingface.co25 days ago81 ptsView details

  2. ornith-ai/Ornith-1.5-397B

    text-generation · transformers · safetensors · qwen3_5_moe

    huggingface.co26 days ago80 ptsView details

  3. Gautam6/GAUTAMAI

    license:apache-2.0 · region:us

    huggingface.co26 days agoView details

  4. Prannesshkva/Ael-504M

    text-generation · transformers · safetensors · ael

    huggingface.co26 days agoView details

  1. ggml-org/llama.cpp b10590

    <details open> vendor : update subprocess.h (#27409) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/42402532> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/down…

    github.com26 days agoView details

  2. ggml-org/llama.cpp b10589

    <details open> cuda : add POOL_1D support (#27573) * cuda : add POOL_1D support * fix: add missing trailing newline for editorconfig compliance </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/42401257> **macOS…

    github.com26 days agoView details