Skip to content

Archive / 2026-08-18

August 18, 2026

  1. Beware Management Consultants

    about.iceland.co.uk1 month ago622 ptsView detailsJoin discussion

  2. Cerebras CS-4

    cerebras.ai1 month ago457 ptsView detailsJoin discussion

  3. GLM-5.3 Artificial Analysis Benchmarks

    artificialanalysis.ai1 month ago152 ptsView detailsJoin discussion

  4. What If America Went Dark?

    nytimes.com1 month ago39 ptsView detailsJoin discussion

  5. The Microsoft Rebrand Registry

    msrebrandregistry.com1 month ago19 ptsView detailsJoin discussion

  6. The Capital Cycle Theory

    s-1.vercel.app1 month ago18 ptsView detailsJoin discussion

  7. Elon Musk made flying worse so Palantir could profit

    On August 6th, the Minneapolis Air Route Traffic Control Center lost radar and communications for around two hours. The outage disrupted more than 1,100 flights across the center's 330,000 square mile, nine-state airspace sector. Two days earlier, on August 4th, President Donald Trump departed the White House inside h…

    theverge.com1 month ago17 ptsView detailsJoin discussion

  8. Primary source

    ChatGPT Ads expands across Europe

    ChatGPT Ads is expanding to 31 European markets. Learn how advertisers can reach people as they explore, compare options, and make decisions.

    openai.com1 month agoView details

  9. Primary source

    Strengthening democratic oversight in national security

    OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expertise.

    openai.com1 month agoView details

  10. Primary source

    Introducing ChatGPT for Teens: Built for learning, backed by protections

    ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents.

    openai.com1 month agoView details

  11. Primary source

    Pacing model development in an era of cyber-critical capabilities

    OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

    openai.com1 month agoView details

  12. Primary source

    Partnering with CodeAI to prepare the first AI generation

    OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.

    openai.com1 month agoView details

  13. Primary source

    Asana cleared 5 years of engineering work in 2 weeks with Codex

    Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.

    openai.com1 month agoView details

  14. Primary source

    How Much Memory Does Your Agent Actually Need?

    huggingface.co1 month agoView details

  15. NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands

    NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an Apache-2.0 project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands, with no intermediate ONNX export. The build emits a versioned .bundle artifact that runs through native C++ task APIs, so…

    marktechpost.com1 month agoView details

  16. Meet SAM (Sovereign Agent Mesh): A Zero-Config, Zero-Trust P2P Network for AI Agents

    Google has open-sourced SAM (Sovereign Agent Mesh) under Apache-2.0 — and it has nothing to do with Segment Anything. SAM is a zero-config, zero-trust P2P overlay that lets autonomous agents discover and call each other's MCP tools across cloud, on-prem, laptop and edge environments, without exposing a single internal…

    marktechpost.com1 month agoView details

  17. Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

    Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to…

    marktechpost.com1 month agoView details

  18. Robin Williams’ Instagram account brought back to fight ‘AI abuse’

    Robin Williams' children are taking over their father's Instagram account after his daughter spoke out against the use of his AI likeness, as reported earlier by The Wrap. In a post on Tuesday, Zak, Zelda, and Cody Williams write that they want the late actor's Instagram profile to be a "safe, trusted place where the…

    theverge.com1 month agoView details

  19. OpenAI lays out new security changes after its AI hacked Hugging Face

    OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks co…

    theverge.com1 month agoView details

  20. OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

    The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.

    wired.com1 month agoView details

  21. Firefox’s Smart Window promises a better AI browser

    Starting today, AI chats in Firefox's Smart Window AI browsing mode can pull from current web info and show source links in chat responses through a partnership with Exa. Smart Window can also now automatically suggest tab groups and show visual previews of pages you previously visited when you search your browsing hi…

    theverge.com1 month agoView details

  22. Google’s Pet Memory forgot who my cats are

    One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, I thought th…

    theverge.com1 month agoView details

  23. ChatGPT is getting a dedicated mode for teens

    ChatGPT for Teens includes safeguards and parental controls. | Image: OpenAI OpenAI is introducing a dedicated ChatGPT mode for teenagers, combining existing youth safeguards and new safety features under one roof. The launch comes amid mounting public scrutiny over how AI tools affect younger users, as other platform…

    theverge.com1 month agoView details

  24. Apple’s camera-equipped AirPods appear in leaked video

    The AirPods in the video look like a chunkier version of Apple’s AirPods Pro 3. | Image: Apple / MacRumors We may have our first glimpse of Apple's rumored camera-equipped AirPods, thanks to a video that MacRumors found in the macOS Tahoe 26.7 Release Candidate. The short video clip features a man - who is wearing the…

    theverge.com1 month agoView details

  25. The Powerful Chinese AI Model Experts Warned About Is Here

    Z.ai’s latest AI model release could help companies secure their systems—or find its way into the hands of hackers.

    wired.com1 month agoView details

  26. J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

    arXiv:2608.17063v1 Announce Type: cross Abstract: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the mod…

    arxiv.org1 month agoView details

  27. When AI Designs AI: Innovation or Imitation?

    arXiv:2608.17471v1 Announce Type: new Abstract: Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic desi…

    arxiv.org1 month agoView details

  28. When to Review: Spaced Repetition for Continual Pre-Training of Language Models

    arXiv:2608.17530v1 Announce Type: new Abstract: Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate c…

    arxiv.org1 month agoView details

  29. TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

    arXiv:2608.17588v1 Announce Type: new Abstract: Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate sol…

    arxiv.org1 month agoView details

  30. CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

    arXiv:2608.17911v1 Announce Type: new Abstract: As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability problem in long-term conversational memo…

    arxiv.org1 month agoView details

  31. LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

    arXiv:2608.17330v1 Announce Type: new Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, mul…

    arxiv.org1 month agoView details

  32. Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch

    arXiv:2608.17684v1 Announce Type: new Abstract: Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether learned behavior preserves previously correct behavior or security. We audit SkillOpt, Agent Workflow Memory (AWM), and ReasoningBan…

    arxiv.org1 month agoView details

  33. SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

    arXiv:2608.17468v1 Announce Type: new Abstract: Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying direct…

    arxiv.org1 month agoView details

  34. ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

    arXiv:2608.17856v1 Announce Type: new Abstract: Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Mo…

    arxiv.org1 month agoView details

  35. Children, but not language models, show accelerating returns in word learning

    arXiv:2608.17120v1 Announce Type: new Abstract: Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerati…

    arxiv.org1 month agoView details

  36. TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

    arXiv:2608.17336v1 Announce Type: new Abstract: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial…

    arxiv.org1 month agoView details

  37. KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

    arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-ag…

    arxiv.org1 month agoView details

  38. Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

    arXiv:2608.17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defen…

    arxiv.org1 month agoView details

  39. Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

    arXiv:2608.17499v1 Announce Type: new Abstract: User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repai…

    arxiv.org1 month agoView details

  40. The politics of postmortem privacy

    arXiv:2608.16905v1 Announce Type: cross Abstract: While the existence of postmortem privacy is increasingly acknowledged (such as the protection of the presence of deceased within digital spaces), far less attention has been paid to its internal instability: its scope (the extent of its application), justificatory fou…

    arxiv.org1 month agoView details

  41. The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

    arXiv:2608.17749v1 Announce Type: new Abstract: Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat…

    arxiv.org1 month agoView details

  42. Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

    arXiv:2608.17634v1 Announce Type: new Abstract: The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only th…

    arxiv.org1 month agoView details

  43. Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

    arXiv:2608.17168v1 Announce Type: new Abstract: Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of leg…

    arxiv.org1 month agoView details

  44. Evaluating the Diversity of AI-Generated Content with Diversity Profiles

    arXiv:2608.17731v1 Announce Type: new Abstract: Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarit…

    arxiv.org1 month agoView details

  45. ArborMem: Navigating Interaction States with Memory Forests

    arXiv:2608.17534v1 Announce Type: new Abstract: Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions. Existing methods improve access to conversational history through long-context processing,…

    arxiv.org1 month agoView details

  46. Token Optimization and Context Window Management in Multi-Agent AI Workflows

    arXiv:2608.17188v1 Announce Type: new Abstract: Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents a practitioner framework for token optimization and context-window management, grounded in an internal production dashboard that ext…

    arxiv.org1 month agoView details

  47. An Investigation of the NeurIPS and ICML 2025 Position Tracks

    arXiv:2608.16894v1 Announce Type: cross Abstract: ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper ar…

    arxiv.org1 month agoView details

  48. AutoResearch: Insight In, Hallucination Out

    arXiv:2608.17906v1 Announce Type: new Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Id…

    arxiv.org1 month agoView details

  49. EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

    arXiv:2608.17933v1 Announce Type: new Abstract: Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend he…

    arxiv.org1 month agoView details

  50. Polaris: Learning to Generate Table Descriptions from Retrieval Feedback

    arXiv:2608.17171v1 Announce Type: new Abstract: Many table-centric NLP tasks such as NL2SQL first retrieve relevant tables from large collections using keyword search. Recent work uses LLMs to generate natural-language table descriptions to improve retrieval, but they are typically optimized for fluency rather than re…

    arxiv.org1 month agoView details

  51. Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

    arXiv:2608.17638v1 Announce Type: new Abstract: What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own r…

    arxiv.org1 month agoView details

  52. LLM-Derived Preference Judgments Are Not Self-Consistent

    arXiv:2608.17644v1 Announce Type: new Abstract: Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments…

    arxiv.org1 month agoView details

  53. Chain-of-Experience for Continual LLM Improvement

    arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we re…

    arxiv.org1 month agoView details

  54. Accuracy and Robustness of Model Cascades Under Data Perturbations

    arXiv:2608.17711v1 Announce Type: new Abstract: Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a la…

    arxiv.org1 month agoView details

  55. StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

    arXiv:2608.17800v1 Announce Type: new Abstract: Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that r…

    arxiv.org1 month agoView details

  56. Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks

    arXiv:2608.17352v1 Announce Type: new Abstract: Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Gra…

    arxiv.org1 month agoView details

  57. Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

    arXiv:2608.17583v1 Announce Type: new Abstract: Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing…

    arxiv.org1 month agoView details

  58. Cross-Model Memory Transfer via Target-Side Reader Adaptation

    arXiv:2608.17050v1 Announce Type: new Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric a…

    arxiv.org1 month agoView details

  59. CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

    arXiv:2608.17536v1 Announce Type: new Abstract: Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in…

    arxiv.org1 month agoView details

  60. Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

    arXiv:2608.17950v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true…

    arxiv.org1 month agoView details

  61. Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

    arXiv:2608.17994v1 Announce Type: new Abstract: Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a…

    arxiv.org1 month agoView details

  62. Adaptive Policy Portfolios for Robust Markov Decision Processes

    arXiv:2608.17929v1 Announce Type: new Abstract: Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memo…

    arxiv.org1 month agoView details

  63. KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

    arXiv:2608.17150v1 Announce Type: cross Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not…

    arxiv.org1 month agoView details

  64. AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction

    arXiv:2608.17184v1 Announce Type: new Abstract: Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (L…

    arxiv.org1 month agoView details

  65. The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

    arXiv:2608.18011v1 Announce Type: new Abstract: Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competitio…

    arxiv.org1 month agoView details

  66. The Plot Thins: Uniformity and Linearity in Literary Summaries

    arXiv:2608.17218v1 Announce Type: new Abstract: Works of literature are complicated; they balance plot, suspense, surprise, and artistic expression. Summaries of literature prioritize plot, and therefore may deviate from their sources. Using a combination of manual and LLM-based annotation, we construct a dataset mapp…

    arxiv.org1 month agoView details

  67. GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

    arXiv:2608.17665v1 Announce Type: new Abstract: LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or co…

    arxiv.org1 month agoView details

  68. Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

    arXiv:2608.17153v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. No…

    arxiv.org1 month agoView details

  69. Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

    arXiv:2608.17947v1 Announce Type: new Abstract: Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Pytho…

    arxiv.org1 month agoView details

  70. Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

    arXiv:2608.16975v1 Announce Type: new Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. This amb…

    arxiv.org1 month agoView details

  71. D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

    arXiv:2608.17756v1 Announce Type: new Abstract: Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluati…

    arxiv.org1 month agoView details

  72. The Problem Is the Problem: Towards Scalable Mathematical Discovery

    arXiv:2608.16977v1 Announce Type: new Abstract: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore centra…

    arxiv.org1 month agoView details

  73. Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

    arXiv:2608.17718v1 Announce Type: new Abstract: Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized…

    arxiv.org1 month agoView details

  74. There is No Theoretical Curse of Multilinguality For Embedding Space Structure

    arXiv:2608.17088v1 Announce Type: new Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model…

    arxiv.org1 month agoView details

  75. Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

    arXiv:2608.17223v1 Announce Type: new Abstract: Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a 49,799-article corpus across 16 feature-model…

    arxiv.org1 month agoView details

  76. SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

    arXiv:2608.17501v1 Announce Type: new Abstract: Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI…

    arxiv.org1 month agoView details

  77. DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

    arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate u…

    arxiv.org1 month agoView details

  78. Which Source Wins? Task-Dependent Reliance in Vision-Language Models

    arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the te…

    arxiv.org1 month agoView details

  79. Agent Lightning v1.0: Towards Harnessed Agentic RL

    arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through a…

    arxiv.org1 month agoView details

  80. Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

    arXiv:2608.17051v1 Announce Type: new Abstract: Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We a…

    arxiv.org1 month agoView details

  81. Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention

    arXiv:2608.17288v1 Announce Type: new Abstract: GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model whether strong token features should reinforce or suppress one another. We introduce Q-Interference, a f…

    arxiv.org1 month agoView details

  82. MoNe: Modular Neural Memory for Efficient Long Context Inference

    arXiv:2608.17616v1 Announce Type: new Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-…

    arxiv.org1 month agoView details

  83. Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

    arXiv:2608.17781v1 Announce Type: new Abstract: ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can…

    arxiv.org1 month agoView details

  84. SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

    arXiv:2608.17931v1 Announce Type: new Abstract: Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world app…

    arxiv.org1 month agoView details

  85. Grading Needs a Rubric, Not Intelligence

    arXiv:2608.17938v1 Announce Type: new Abstract: Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at…

    arxiv.org1 month agoView details

  86. Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

    arXiv:2608.18041v1 Announce Type: new Abstract: Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention weights refine that count, which sums every writer in the corpus together. This paper claims a second parameter, phase…

    arxiv.org1 month agoView details

  87. Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

    arXiv:2608.17443v1 Announce Type: new Abstract: Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs)…

    arxiv.org1 month agoView details

  88. SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

    arXiv:2608.17007v1 Announce Type: new Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire in…

    arxiv.org1 month agoView details

  89. Towards Zero-Shot Task Transfer with Neurosymbolic World Models

    arXiv:2608.17959v1 Announce Type: new Abstract: State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-depend…

    arxiv.org1 month agoView details

  90. Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

    arXiv:2608.17625v1 Announce Type: new Abstract: Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms. We deliv…

    arxiv.org1 month agoView details

  91. A decodability criterion predicts when hidden-state selection beats majority voting in large language models

    arXiv:2608.17124v1 Announce Type: new Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so th…

    arxiv.org1 month agoView details

  92. Toward Personal Intelligence Through Cooperative Observation

    arXiv:2608.17128v1 Announce Type: new Abstract: A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a boun…

    arxiv.org1 month agoView details

  93. GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents

    arXiv:2608.16890v1 Announce Type: new Abstract: Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontie…

    arxiv.org1 month agoView details

  94. Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

    arXiv:2608.16891v1 Announce Type: new Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it…

    arxiv.org1 month agoView details

  95. Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

    arXiv:2608.17075v1 Announce Type: new Abstract: Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correc…

    arxiv.org1 month agoView details

  96. A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

    arXiv:2608.17096v1 Announce Type: new Abstract: The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with…

    arxiv.org1 month agoView details

  97. Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection

    arXiv:2608.17170v1 Announce Type: new Abstract: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an autom…

    arxiv.org1 month agoView details

  98. LLM-Only PDDL Domain Repair with Open-Weight Models

    arXiv:2608.17341v1 Announce Type: new Abstract: AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such model…

    arxiv.org1 month agoView details

  99. Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits

    arXiv:2608.17741v1 Announce Type: new Abstract: OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entai…

    arxiv.org1 month agoView details

  100. Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

    arXiv:2608.17687v1 Announce Type: new Abstract: Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is…

    arxiv.org1 month agoView details

  101. FedPref: Federated Preference Learning for Structured Radiology Report Extraction

    arXiv:2608.16971v1 Announce Type: new Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evi…

    arxiv.org1 month agoView details

  102. Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

    arXiv:2608.17102v1 Announce Type: new Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared af…

    arxiv.org1 month agoView details

  103. Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

    arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large…

    arxiv.org1 month agoView details

  104. Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

    arXiv:2608.17247v1 Announce Type: new Abstract: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, the…

    arxiv.org1 month agoView details

  105. Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

    arXiv:2608.17270v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similar…

    arxiv.org1 month agoView details

  106. Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

    arXiv:2608.17574v1 Announce Type: new Abstract: How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein…

    arxiv.org1 month agoView details

  107. Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

    arXiv:2608.17744v1 Announce Type: new Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only…

    arxiv.org1 month agoView details

  108. The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

    arXiv:2608.16956v1 Announce Type: cross Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered…

    arxiv.org1 month agoView details

  109. ASI-Bench: At the Dawn of Artificial Superintelligence

    arXiv:2608.17271v1 Announce Type: new Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on lear…

    arxiv.org1 month agoView details

  110. DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation

    arXiv:2608.17282v1 Announce Type: new Abstract: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework tha…

    arxiv.org1 month agoView details

  111. PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

    arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories d…

    arxiv.org1 month agoView details

  112. LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

    arXiv:2608.17299v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks pr…

    arxiv.org1 month agoView details

  113. SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

    arXiv:2608.17301v1 Announce Type: new Abstract: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems rema…

    arxiv.org1 month agoView details

  114. Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

    arXiv:2608.17319v1 Announce Type: new Abstract: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignme…

    arxiv.org1 month agoView details

  115. ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation

    arXiv:2608.17356v1 Announce Type: new Abstract: Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three de…

    arxiv.org1 month agoView details

  116. PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

    arXiv:2608.17379v1 Announce Type: new Abstract: We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over f…

    arxiv.org1 month agoView details

  117. An Investigation of Translationese in the Generations of Multilingual Large Language Models

    arXiv:2608.17399v1 Announce Type: new Abstract: Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, i…

    arxiv.org1 month agoView details

  118. From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis

    arXiv:2608.17454v1 Announce Type: new Abstract: This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level cov…

    arxiv.org1 month agoView details

  119. Effects of Answer Format Variation on Gender Bias in Large Language Models

    arXiv:2608.17516v1 Announce Type: new Abstract: Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a…

    arxiv.org1 month agoView details

  120. Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

    arXiv:2608.17587v1 Announce Type: new Abstract: Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inferen…

    arxiv.org1 month agoView details

  121. Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

    arXiv:2608.17605v1 Announce Type: new Abstract: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context acro…

    arxiv.org1 month agoView details

  122. TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

    arXiv:2608.17795v1 Announce Type: new Abstract: Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database c…

    arxiv.org1 month agoView details

  123. Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

    arXiv:2608.17810v1 Announce Type: new Abstract: The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM perfor…

    arxiv.org1 month agoView details

  124. From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

    arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this pape…

    arxiv.org1 month agoView details

  125. Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

    arXiv:2608.17843v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a con…

    arxiv.org1 month agoView details

  126. BayesPrompt: human readable prompts that make sense

    arXiv:2608.17866v1 Announce Type: new Abstract: Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the perplexity of a given answer, however, consistently yield so-called pseudoprompts, unintelligible stri…

    arxiv.org1 month agoView details

  127. BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

    arXiv:2608.17895v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external d…

    arxiv.org1 month agoView details

  128. TokEval: A Tokenizer Evaluation Suite

    arXiv:2608.18062v1 Announce Type: new Abstract: Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstr…

    arxiv.org1 month agoView details

  129. Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

    arXiv:2608.18072v1 Announce Type: new Abstract: Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictate…

    arxiv.org1 month agoView details

  130. Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis

    arXiv:2311.06273v1 Announce Type: cross Abstract: The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Recognizing its value across different areas, our study investigates ChatGPT's capacity to predict stock market movements using only…

    arxiv.org1 month agoView details

  131. Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

    arXiv:2602.14784v1 Announce Type: cross Abstract: Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return…

    arxiv.org1 month agoView details

  132. Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

    arXiv:2608.15382v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-c…

    arxiv.org1 month agoView details

  133. When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

    arXiv:2608.16909v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok…

    arxiv.org1 month agoView details

  134. SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback

    arXiv:2608.16934v1 Announce Type: cross Abstract: RTL code generation is a critical stage in hardware design, and the emergence of agentic systems offers new opportunities to automate this process. To generate correct RTL code, agents must understand sequential behavior, including how signals evolve and propagate over…

    arxiv.org1 month agoView details

  135. Memory Is Communication: The Frontier Between Remembering and Signaling

    arXiv:2608.17053v1 Announce Type: cross Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an a…

    arxiv.org1 month agoView details

  136. Uncertainty-Aware Decision Making in Multimodal Large Language Models

    arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input qualit…

    arxiv.org1 month agoView details

  137. What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?

    arXiv:2608.17325v1 Announce Type: new Abstract: Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across languages. In this work, we analyze what tokens are learned when tokenization is jointly optim…

    arxiv.org1 month agoView details

  138. Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

    arXiv:2608.17809v1 Announce Type: new Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLM…

    arxiv.org1 month agoView details

  139. When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

    arXiv:2608.17979v1 Announce Type: new Abstract: Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI-assisted writ…

    arxiv.org1 month agoView details

  140. LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

    arXiv:2608.17393v1 Announce Type: new Abstract: Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradi…

    arxiv.org1 month agoView details

  141. Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations

    arXiv:2608.17433v1 Announce Type: new Abstract: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the sam…

    arxiv.org1 month agoView details

  142. Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression

    arXiv:2608.17434v1 Announce Type: new Abstract: We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variation budget A, and output bound B. For this O(L w^2)-parameterized architecture, the known lower and upper bounds differ by one factor…

    arxiv.org1 month agoView details

  1. orcarouter/Qwen3.8-27B-Uncensored-GGUF

    image-text-to-text · gguf · abliterated · qwen

    huggingface.co1 month ago310 ptsView details

  2. peculiar-ragdoll/Qwen-Sharp-Chat-Templates

    mlx · jinja · chat-template

    huggingface.co1 month ago280 ptsView details

  3. DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF

    image-text-to-text · gguf · qwen3_5 · unsloth

    huggingface.co1 month ago272 ptsView details

  4. z-lab/Qwen3.8-27B-DFlash2

    text-generation · transformers · safetensors · qwen3

    huggingface.co1 month ago241 ptsView details

  5. LBH-123-AI/Minimax_h3_latent_Upscaler

    region:us

    huggingface.co1 month ago212 ptsView details

  6. incoai/Qwen3.8-27B-DFlash2

    text-generation · transformers · safetensors · qwen3

    huggingface.co1 month ago181 ptsView details

  7. incoai/Qwen3.8-27B-DFlash2-GGUF

    text-generation · llama.cpp · gguf · dflash2

    huggingface.co1 month ago121 ptsView details

  8. z-lab/Qwen3.8-27B-DFlash2-GGUF

    text-generation · llama.cpp · gguf · dflash2

    huggingface.co1 month ago85 ptsView details

  9. esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF

    text-generation · gguf · nvfp4 · qwen3.8

    huggingface.co1 month ago54 ptsView details

  10. mlasli/Qwen3.8-27B-Heretic-Uncensored-Q4_K_M-GGUF

    text-generation · gguf · qwen3.8 · qwen3

    huggingface.co1 month ago2 ptsView details

  11. nmthien/vietnamese-gpt2

    text-generation · transformers · safetensors · gpt2

    huggingface.co1 month ago1 ptsView details

  12. openeurollm/oellm-9b-256k-sft

    text-generation · transformers · safetensors · qwen3

    huggingface.co1 month ago1 ptsView details

  13. stage-babylm/llama-384-2L

    text-generation · transformers · safetensors · llama

    huggingface.co1 month agoView details

  14. MohamedAhmedAE/llava-medical-8B-clip-vit_kaggle-stage2

    safetensors · llava · region:us

    huggingface.co1 month agoView details

  15. sullivan1502/base-action-grpo

    text-generation · transformers · safetensors · llama

    huggingface.co1 month agoView details

  1. openai/openai-python v3.3.0

    ## [3.3.0](https://github.com/openai/openai-python/compare/v3.2.0...v3.3.0) (2026-08-18) ### Features * support named data-residency endpoints ([#3646](https://github.com/openai/openai-python/issues/3646)) ([11ee914](https://github.com/openai/openai-python/commit/11ee91475694d9c…

    github.com1 month agoView details

  2. langchain-ai/langchain langchain-openai==1.5.2

    Changes since langchain-openai==1.5.1 release(openai): 1.5.2 (#39719) fix(openai): preserve reasoning item boundaries (#39278) release(openai): 1.5.2a1 (#39709) feat(openai): extract gateway metadata from response headers when available (#39706) chore(openai): update snapshots (…

    github.com1 month agoView details

  3. modelcontextprotocol/servers 2026.8.18

    # Release : v2026.8.18 ## Updated packages - @modelcontextprotocol/server-everything@2026.8.18 - mcp-server-time@2026.8.18 - mcp-server-fetch@2026.8.18 - mcp-server-git@2026.8.18

    github.com1 month agoView details

  4. ggml-org/llama.cpp b10488

    <details open> ci : Update OpenVINO to 2026.3, skip nemotron-h rollback test (#27292) * update to ov-2026.3, update device drivers * ci: skip nemotron-h rollback test on OpenVINO The OpenVINO backend does not support SSM_SCAN, so the Nemotron-H recurrent state rollback graph is…

    github.com1 month agoView details

  5. ggml-org/llama.cpp b10486

    <details open> mtmd: fix LFM2 image tiling threshold (#27057) * mtmd: fix LFM2 image tiling threshold * refactor testing * fix * fix on windows --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Ap…

    github.com1 month agoView details

  6. ggml-org/llama.cpp v0.1.2

    > [!NOTE] > Semantic versioning is still work in progress. > More info can be found in https://github.com/ggml-org/ggml/discussions/1579 **Nightly build:** [b10485](https://github.com/ggml-org/llama.cpp/releases/tag/b10485) ## Change log since v0.1.1 1511ce3bc sync : ggml da786d…

    github.com1 month agoView details