Skip to content

Archive / 2026-08-12

August 12, 2026

  1. DeepSeek V4 Pro 0813

    openrouter.ai1 month ago1021 ptsView detailsJoin discussion

  2. Grok 4.6

    x.ai1 month ago625 ptsView detailsJoin discussion

  3. What sort of maths are LLMs good at?

    gowers.wordpress.com1 month ago261 ptsView detailsJoin discussion

  4. My Agent Setup

    chad.cm1 month ago132 ptsView detailsJoin discussion

  5. I'm Done Coding with AI

    youtube.com1 month ago39 ptsView detailsJoin discussion

  6. SpaceXAI: Grok 4.6

    openrouter.ai1 month ago31 ptsView detailsJoin discussion

  7. I'm Done Using AI

    brettcodes.com1 month ago19 ptsView detailsJoin discussion

  8. When Will AI Take My Job?

    whenwillaitakemyjob.ai1 month ago17 ptsView detailsJoin discussion

  9. Reporter gets $850000 in raid suit

    marionrecord.com1 month ago14 ptsView detailsJoin discussion

  10. IT Unemployment Rate Jumps to 6.7%

    itmanager.substack.com1 month ago13 ptsView detailsJoin discussion

  11. Primary source

    Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Generative AI

    research.google1 month ago10 ptsView details

  12. Google Pixel Tag

    blog.google1 month ago10 ptsView detailsJoin discussion

  13. A tycoon game where you run an AI company

    store.steampowered.com1 month ago10 ptsView detailsJoin discussion

  14. Primary source

    From assistance to execution: How enterprises put AI to work

    OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.

    openai.com1 month agoView details

  15. Primary source

    Putting sign language AI into users’ hands

    Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

    deepmind.google1 month agoView details

  16. AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation

    Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework. This comprehensive guide walks through Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (GRPO), optimized to run efficiently on 16GB hardware without needing heavy di…

    marktechpost.com1 month agoView details

  17. NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router

    NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.

    marktechpost.com1 month agoView details

  18. Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark

    Object removal models have improved faster than the metrics used to judge them. Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR, SSIM, LPIPS, ReMOVE and CFD frequently rank their outputs the wrong way. The root cause is structural: erasure is an ill-posed, one-to-ma…

    marktechpost.com1 month agoView details

  19. The White House Is Going to Expand Its AI Policy

    Open models may soon be added to an updated AI framework, sources tell WIRED, as the White House continues to grapple with how to regulate a technology it has tried not to regulate.

    wired.com1 month agoView details

  20. Rogue AI Agents Aren’t Evil. They’re Just Eager to Please

    AI agents that break free and hack into other systems are only trying to make us happy.

    wired.com1 month agoView details

  21. Twitch streamers can now opt out from training Amazon’s AI

    Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. Opting out means that "your streams, VODs, clips, stream chats, and pictures and text on your channel" won't be used in "future training" of an Amazon AI model "whose purpose is to generate or synthesize text, aud…

    theverge.com1 month agoView details

  22. Guitar company D’Addario admits that AI music was used in a promotional video

    After weeks of controversy and speculation, music company D'Addario has admitted that AI, specifically Suno, was used as part of a recent promotional video. For nearly two weeks, the company has denied the allegations, even as evidence piled up against it. It offered various explanations, from low-quality exports, to…

    theverge.com1 month agoView details

  23. 4 New Camera Tricks on Google’s Latest Pixel 11 Smartphones

    From Magic Capture and Instant Night Sight to a built-in teleprompter, here’s a look at a few camera features on Google’s new Pixel 11 series.

    wired.com1 month agoView details

  24. Google’s Pixel Watch 5 dives deeper into AI and health

    At least there’s no new proprietary charger this year. Huzzah!! | Photo: David Imel / The Verge The $399 Google Pixel Watch 5 isn't about the hardware. Sure, there's a new satin pyrite case finish, a few new strap colors, and a Steph Curry Special Edition. Under the hood, there's a slightly faster Qualcomm processor a…

    theverge.com1 month agoView details

  25. Of course the ChatGPT dog cancer vaccine spawned a startup

    Remember that much-hyped story about an Australian tech entrepreneur using ChatGPT, Grok, and other AI tools to craft a personalized cancer vaccine for his dog? Well, surprise: He's launched a startup. That entrepreneur is Paul Conyngham, who says he is launching Gamgee to offer "personalised mRNA cancer vaccines for…

    theverge.com1 month agoView details

  26. Grok is now an AI ‘teammate’ you can assign work

    You’ll have to be fine with letting Grok sign into your online accounts, however. | Image: SpaceXAI SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent "AI teammates" that can do your work for you. The bots share their own cloud-based computer environment, and can sign i…

    theverge.com1 month agoView details

  27. The Job-Interview Tattoo Guy Everyone Got Mad at Finally Explains Himself

    LemonLime cofounder Jordan Zietz hears your criticism loud and clear. That’s why he got his startup’s logo tattooed on his shoulder.

    wired.com1 month agoView details

  28. Oh Lord, AI Reporters Are Actually Breaking Big News

    Last week, an AI newsroom beat mainstream journalists—including WIRED—to a story about OpenAI and hacking. It’s just the beginning.

    wired.com1 month agoView details

  29. You’re Thinking About Online Trends All Wrong

    From pessimism around dating to AI reshaping culture, cyber-ethnographer Ruby J. Thelot tells WIRED why people are putting too much stock into things that go viral.

    wired.com1 month agoView details

  30. Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

    arXiv:2608.11528v1 Announce Type: new Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and o…

    arxiv.org1 month agoView details

  31. CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

    arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrenc…

    arxiv.org1 month agoView details

  32. Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

    arXiv:2608.11552v1 Announce Type: new Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clari…

    arxiv.org1 month agoView details

  33. DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

    arXiv:2608.11441v1 Announce Type: new Abstract: Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due t…

    arxiv.org1 month agoView details

  34. ODE-Based Transformer Decoders for Iterative Sign Language Translation

    arXiv:2608.11352v1 Announce Type: new Abstract: Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. We propose a parameter-efficient alternative that improves expressiveness without in…

    arxiv.org1 month agoView details

  35. CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

    arXiv:2608.11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In c…

    arxiv.org1 month agoView details

  36. On Weak Bisimilarities in CCSK

    arXiv:2608.11531v1 Announce Type: new Abstract: In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences and commonalities. In particular, for the weak reversible case, not previously studied in the literatu…

    arxiv.org1 month agoView details

  37. LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs

    arXiv:2608.11220v1 Announce Type: new Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually. Applying artificial intelligence in the task could potentially lead not only to process automati…

    arxiv.org1 month agoView details

  38. When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

    arXiv:2608.11715v1 Announce Type: new Abstract: The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Alt…

    arxiv.org1 month agoView details

  39. Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

    arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that ou…

    arxiv.org1 month agoView details

  40. Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

    arXiv:2608.11573v1 Announce Type: new Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-l…

    arxiv.org1 month agoView details

  41. Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library

    arXiv:2608.11767v1 Announce Type: new Abstract: When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision…

    arxiv.org1 month agoView details

  42. The Edge-based Contiguous p-median Problem with Connections to Logistics Districting

    arXiv:2608.11230v1 Announce Type: new Abstract: This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact and contiguous territories. Two binary programming models are introduced, both of which incorporate a network distance. The first mo…

    arxiv.org1 month agoView details

  43. Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

    arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a co…

    arxiv.org1 month agoView details

  44. Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

    arXiv:2608.11215v1 Announce Type: cross Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn…

    arxiv.org1 month agoView details

  45. Proportional Analogies on Probability Distributions via Bayesian Updating

    arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postulates. While proport…

    arxiv.org1 month agoView details

  46. Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models

    arXiv:2608.11657v1 Announce Type: new Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishing a non-linear homeostatic feedback loop…

    arxiv.org1 month agoView details

  47. Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

    arXiv:2608.11742v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only…

    arxiv.org1 month agoView details

  48. RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

    arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming fe…

    arxiv.org1 month agoView details

  49. LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification

    arXiv:2608.11753v1 Announce Type: new Abstract: Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as an additional input on the task of classifying sentences from Federal Reserve com…

    arxiv.org1 month agoView details

  50. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

    arXiv:2608.12062v1 Announce Type: new Abstract: Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed t…

    arxiv.org1 month agoView details

  51. Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

    arXiv:2608.11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these system…

    arxiv.org1 month agoView details

  52. Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

    arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates witho…

    arxiv.org1 month agoView details

  53. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

    arXiv:2608.11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setti…

    arxiv.org1 month agoView details

  54. Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

    arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literat…

    arxiv.org1 month agoView details

  55. A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph

    arXiv:2608.11211v1 Announce Type: new Abstract: Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our verifiable contribu…

    arxiv.org1 month agoView details

  56. Harnessing agent memory to build lifelong AI partners for materials scientists

    arXiv:2608.11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential for reproducibility…

    arxiv.org1 month agoView details

  57. Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones

    arXiv:2608.11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. We distinguish thre…

    arxiv.org1 month agoView details

  58. Stigma and Support in Online Sexual Violence Narratives on Reddit

    arXiv:2608.11433v1 Announce Type: new Abstract: Online communities increasingly provide spaces where survivors of sexual violence can share their experiences and seek support. Although prior research has examined stigma and social support separately, less is known about how stigma expressed in survivor narratives rela…

    arxiv.org1 month agoView details

  59. Hybrid Gated Attention

    arXiv:2608.11805v1 Announce Type: new Abstract: Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of…

    arxiv.org1 month agoView details

  60. Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

    arXiv:2608.11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal i…

    arxiv.org1 month agoView details

  61. Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

    arXiv:2608.11249v1 Announce Type: new Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compress…

    arxiv.org1 month agoView details

  62. Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

    arXiv:2608.12218v2 Announce Type: new Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to ric…

    arxiv.org1 month agoView details

  63. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

    arXiv:2608.11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging…

    arxiv.org1 month agoView details

  64. Locating and Controlling Implicit Personalization in Large Language Models

    arXiv:2608.11735v1 Announce Type: new Abstract: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal ac…

    arxiv.org1 month agoView details

  65. Easper: An Accessible ASR Pipeline for Language Documentation

    arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflo…

    arxiv.org1 month agoView details

  66. A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement

    arXiv:2608.12269v1 Announce Type: new Abstract: Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, stakeholder comments and publicly available government data are often underutilized,…

    arxiv.org1 month agoView details

  67. TELLME: Test-Enhanced Learning for Language Model Enrichment

    arXiv:2608.11788v1 Announce Type: new Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational…

    arxiv.org1 month agoView details

  68. Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

    arXiv:2608.11822v1 Announce Type: new Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeli…

    arxiv.org1 month agoView details

  69. HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry

    arXiv:2608.11768v1 Announce Type: new Abstract: The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally construct rule antecedents…

    arxiv.org1 month agoView details

  70. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

    arXiv:2608.11616v2 Announce Type: new Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the…

    arxiv.org1 month agoView details

  71. Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

    arXiv:2608.11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer…

    arxiv.org1 month agoView details

  72. Benchmarking LLM Judges for Mobile Agent Evaluation

    arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge m…

    arxiv.org1 month agoView details

  73. Learning from Online User Feedback for Shopping Agents

    arXiv:2608.11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offli…

    arxiv.org1 month agoView details

  74. Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

    arXiv:2608.12278v1 Announce Type: new Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation bench…

    arxiv.org1 month agoView details

  75. Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

    arXiv:2608.11212v1 Announce Type: cross Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes…

    arxiv.org1 month agoView details

  76. Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

    arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only \emph{revealed} or magn…

    arxiv.org1 month agoView details

  77. The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

    arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the m…

    arxiv.org1 month agoView details

  78. InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

    arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can hand…

    arxiv.org1 month agoView details

  79. HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

    arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understandi…

    arxiv.org1 month agoView details

  80. AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection

    arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thoro…

    arxiv.org1 month agoView details

  81. MaSRead: Content-Addressed Reading of Replicated Latent Stores

    arXiv:2608.11218v1 Announce Type: new Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query,…

    arxiv.org1 month agoView details

  82. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

    arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries…

    arxiv.org1 month agoView details

  83. EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

    arXiv:2608.11584v1 Announce Type: new Abstract: Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple quer…

    arxiv.org1 month agoView details

  84. From Monolithic to Modular: Segment-level Automatic Prompt Optimization

    arXiv:2608.11219v1 Announce Type: cross Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted imp…

    arxiv.org1 month agoView details

  85. VQ-bench: A Composable Vector Quantization Framework

    arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research activity. This paper provides a unified framework for developing and benchmarking new quantization algorit…

    arxiv.org1 month agoView details

  86. TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

    arXiv:2608.11236v1 Announce Type: new Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decompos…

    arxiv.org1 month agoView details

  87. Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

    arXiv:2608.08514v1 Announce Type: cross Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models). The first, RPC, aggregates token p…

    arxiv.org1 month agoView details

  88. Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

    arXiv:2608.11233v1 Announce Type: new Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-pres…

    arxiv.org1 month agoView details

  89. Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

    arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware i…

    arxiv.org1 month agoView details

  90. Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction

    arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction remains challenging: local errors accumulate as spectral inconsistency, phase misalignment, or mean drif…

    arxiv.org1 month agoView details

  91. BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model

    arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, multimodal knowledge…

    arxiv.org1 month agoView details

  92. Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

    arXiv:2608.11245v1 Announce Type: new Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learni…

    arxiv.org1 month agoView details

  93. Towards the Harness of Embodied Agents

    arXiv:2608.11246v1 Announce Type: new Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harne…

    arxiv.org1 month agoView details

  94. AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

    arXiv:2608.11250v1 Announce Type: new Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen…

    arxiv.org1 month agoView details

  95. CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

    arXiv:2608.11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existi…

    arxiv.org1 month agoView details

  96. AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention

    arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task-specific know…

    arxiv.org1 month agoView details

  97. GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

    arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are no…

    arxiv.org1 month agoView details

  98. XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

    arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's inte…

    arxiv.org1 month agoView details

  99. FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

    arXiv:2608.11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while…

    arxiv.org1 month agoView details

  100. When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

    arXiv:2608.11843v1 Announce Type: new Abstract: The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the historical fact rather than merely the surface. Low-resource historical…

    arxiv.org1 month agoView details

  101. Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

    arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition…

    arxiv.org1 month agoView details

  102. Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

    arXiv:2608.11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Em…

    arxiv.org1 month agoView details

  103. From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

    arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a front…

    arxiv.org1 month agoView details

  104. When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

    arXiv:2608.11403v1 Announce Type: new Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy…

    arxiv.org1 month agoView details

  105. A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

    arXiv:2608.11483v1 Announce Type: new Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source f…

    arxiv.org1 month agoView details

  106. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

    arXiv:2608.12121v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by t…

    arxiv.org1 month agoView details

  107. CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

    arXiv:2608.11588v1 Announce Type: new Abstract: Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework tha…

    arxiv.org1 month agoView details

  108. Foresight Without Seeing: Latent Futures for World Action Models

    arXiv:2608.11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Explicit-future WAMs…

    arxiv.org1 month agoView details

  109. Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects

    arXiv:2608.12018v1 Announce Type: new Abstract: Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that diverge significantly from Standard Colloquial Bangla (SCB) in phonology, morpholo…

    arxiv.org1 month agoView details

  110. One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

    arXiv:2608.12253v1 Announce Type: new Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM…

    arxiv.org1 month agoView details

  111. Adaptive Hybrid Particle Swarm Optimization with Gradient Descent

    arXiv:2608.11258v1 Announce Type: new Abstract: Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not universally. We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function on swarm diversity to automatically modulate gradient…

    arxiv.org1 month agoView details

  112. Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

    arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem0, Hindsight, and Mastra Observa…

    arxiv.org1 month agoView details

  113. Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

    arXiv:2608.11924v1 Announce Type: new Abstract: Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generatio…

    arxiv.org1 month agoView details

  114. SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

    arXiv:2608.12129v1 Announce Type: new Abstract: While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this…

    arxiv.org1 month agoView details

  115. Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

    arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harn…

    arxiv.org1 month agoView details

  116. Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

    arXiv:2608.11660v1 Announce Type: new Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledg…

    arxiv.org1 month agoView details

  117. Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

    arXiv:2608.11354v1 Announce Type: new Abstract: Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive ext…

    arxiv.org1 month agoView details

  118. LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

    arXiv:2608.11919v1 Announce Type: new Abstract: Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor…

    arxiv.org1 month agoView details

  119. LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

    arXiv:2608.11922v1 Announce Type: new Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.4769…

    arxiv.org1 month agoView details

  120. Forecasting Side Effects of Activation Steering

    arXiv:2608.11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to dep…

    arxiv.org1 month agoView details

  121. Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

    arXiv:2608.12149v1 Announce Type: new Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist throug…

    arxiv.org1 month agoView details

  122. A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems

    arXiv:2608.11221v1 Announce Type: new Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these systems emerges from the interaction between those artefacts and their operational environment. Sim…

    arxiv.org1 month agoView details

  123. Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures

    arXiv:2608.11255v1 Announce Type: new Abstract: Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. While deep learning models can provide accurate predic…

    arxiv.org1 month agoView details

  124. From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

    arXiv:2608.11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empiric…

    arxiv.org1 month agoView details

  125. Making AI-Generated Feedback Matter: From Provision to Student Enactment

    arXiv:2608.11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedba…

    arxiv.org1 month agoView details

  126. Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

    arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based m…

    arxiv.org1 month agoView details

  127. Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning

    arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions…

    arxiv.org1 month agoView details

  128. Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

    arXiv:2608.11332v1 Announce Type: new Abstract: Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transc…

    arxiv.org1 month agoView details

  129. Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

    arXiv:2608.11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher a…

    arxiv.org1 month agoView details

  130. Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

    arXiv:2608.11408v1 Announce Type: new Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether t…

    arxiv.org1 month agoView details

  131. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

    arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $\tau^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset…

    arxiv.org1 month agoView details

  132. A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

    arXiv:2608.12138v1 Announce Type: new Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate…

    arxiv.org1 month agoView details

  133. Asymptotic Risk Calibration for Selective Question Answering

    arXiv:2608.12008v1 Announce Type: new Abstract: Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores cannot perfectly distinguish correct predictions from incorrect ones, and directly a…

    arxiv.org1 month agoView details

  134. Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

    arXiv:2608.11242v1 Announce Type: new Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behav…

    arxiv.org1 month agoView details

  135. Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

    arXiv:2608.11338v1 Announce Type: new Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills. Existing works focus on performance gain over cost effectiveness. As a result, li…

    arxiv.org1 month agoView details

  136. Self-Evolving Embodied Agents via Skill-Harness Evolution

    arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement…

    arxiv.org1 month agoView details

  137. Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

    arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically follow two paths: training compact models…

    arxiv.org1 month agoView details

  138. Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

    arXiv:2608.11947v1 Announce Type: new Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option…

    arxiv.org1 month agoView details

  139. EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

    arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time. In particular,…

    arxiv.org1 month agoView details

  140. Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration

    arXiv:2608.11460v1 Announce Type: new Abstract: Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the kinds of prompting traits that contribute to task success. Moreover, we need to unc…

    arxiv.org1 month agoView details

  141. Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

    arXiv:2608.11786v1 Announce Type: new Abstract: Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that…

    arxiv.org1 month agoView details

  142. Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

    arXiv:2608.11649v1 Announce Type: new Abstract: As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that i…

    arxiv.org1 month agoView details

  143. Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

    arXiv:2608.11232v1 Announce Type: new Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A…

    arxiv.org1 month agoView details

  144. Structuring the Space of Perspectives

    arXiv:2608.12113v1 Announce Type: new Abstract: The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization. A wide range of operative co…

    arxiv.org1 month agoView details

  145. Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

    arXiv:2608.11624v1 Announce Type: new Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to h…

    arxiv.org1 month agoView details

  146. The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

    arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routi…

    arxiv.org1 month agoView details

  1. Qwen/Qwen3.8-2.4T-A95B

    text-generation · transformers · safetensors · qwen3_5_moe_text

    huggingface.co1 month ago1164 ptsView details

  2. DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

    image-text-to-text · gguf · MTP GGUFS · Regular GGUFS

    huggingface.co1 month ago373 ptsView details

  3. Gazingstars123/Anima-2.9B

    text-to-image · diffusion-single-file · anima · comfyui

    huggingface.co1 month ago323 ptsView details

  4. Qwen/Qwen3.8-2.4T-A95B-FP8

    text-generation · transformers · safetensors · qwen3_5_moe_text

    huggingface.co1 month ago228 ptsView details

  5. LiquidAI/LFM2.5-VL-3B

    image-text-to-text · transformers · safetensors · lfm2_vl

    huggingface.co1 month ago179 ptsView details

  6. IndexTeam/IndexTTS-2.5

    text-to-speech · indextts · safetensors · text-to-speech

    huggingface.co1 month ago137 ptsView details

  7. CohereLabs/North-Micro-Vision-Instruct

    image-text-to-text · transformers · safetensors · cohere_compass

    huggingface.co1 month ago120 ptsView details

  8. unsloth/Qwen3.8-2.4T-A95B-GGUF

    text-generation · transformers · gguf · unsloth

    huggingface.co1 month ago105 ptsView details

  9. Kutches/Kr3a

    gguf · region:us

    huggingface.co1 month ago60 ptsView details

  10. zenosai/MonkeyOCRv2-B-Parsing

    image-text-to-text · transformers · safetensors · monkeyocrv2

    huggingface.co1 month ago13 ptsView details

  11. Kutches/Anim4

    gguf · region:us

    huggingface.co1 month ago9 ptsView details

  12. ibyteohdear/Lightricks-LTX-2

    image-to-video · diffusion-single-file · image-to-video · text-to-video

    huggingface.co1 month ago4 ptsView details

  13. andreaborio/DeepSeek-V4-Flash-Hebrus-GGUF

    text-generation · hebrus · gguf · deepseek

    huggingface.co1 month ago2 ptsView details

  14. NN-Dataset/tflite

    tflite · region:us

    huggingface.co1 month ago2 ptsView details

  15. ReliquaryForge/qwen3.5-4b-reliquary-v4

    safetensors · qwen3_5 · region:us

    huggingface.co1 month ago2 ptsView details

  16. andreaborio/Qwen3.6-35B-A3B-Hebrus-GGUF

    text-generation · hebrus · gguf · qwen3.6

    huggingface.co1 month ago1 ptsView details

  17. deepdml/whisper-tiny-es-mix-norm

    automatic-speech-recognition · transformers · tensorboard · safetensors

    huggingface.co1 month ago1 ptsView details

  18. Nhoodie/npu-moe-vlm

    region:us

    huggingface.co1 month ago1 ptsView details

  19. MohamedAhmedAE/llava-medical-3B-clip-vit-stage2

    safetensors · llava · region:us

    huggingface.co1 month ago1 ptsView details

  20. njand/wav2vec2-xls-r-latin

    automatic-speech-recognition · onnx · safetensors · wav2vec2

    huggingface.co1 month agoView details

  21. MohamedAhmedAE/llava-medical-1B-clip-vit-stage2

    safetensors · llava · region:us

    huggingface.co1 month agoView details

  22. khairi/life2lang-base-wo-pt-it

    safetensors · t5 · region:us

    huggingface.co1 month agoView details

  1. Carasibana/ComfyUI-H3-FaceRefine

    Refine and improve the quality of small faces in MiniMax H3 video. Per-frame face tracking, crop, refine with H3, stitch back.

    github.com1 month ago26 ptsView details

  2. langchain-ai/langchain langchain-anthropic==1.5.6

    Changes since langchain-anthropic==1.5.5 release(anthropic): 1.5.6 (#39622) fix(anthropic): normalize `tool_search_tool_result` blocks (#39621) fix(anthropic): correct model profile data for Fable 5, Sonnet 5, Opus 4.1 (#39604)

    github.com1 month agoView details

  3. ggml-org/llama.cpp b10375

    <details open> chat : tighten bare function parsing for Qwen models (#26793) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10375/llama-b10375-bin-macos-arm64.tar.gz) - macOS A…

    github.com1 month agoView details

  4. ggml-org/llama.cpp b10373

    <details open> imatrix.cpp: Move finite check and only check touched experts (#26861) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10373/llama-b10373-bin-macos-arm64.tar.gz)…

    github.com1 month agoView details