Archive / 2026-08-18
August 18, 2026
News
View all news →about.iceland.co.uk1 month ago622 ptsView detailsJoin discussion
Google has acquired the data of failed US airline Spirit
theregister.com1 month ago608 ptsView detailsJoin discussion
cerebras.ai1 month ago457 ptsView detailsJoin discussion
Using the railway network as a flatbed scanner
philo.gay1 month ago441 ptsView detailsJoin discussion
Sticky wage norms and the real wage cost of unexpected inflation
bfi.uchicago.edu1 month ago390 ptsView detailsJoin discussion
fx :Tiny, open, native coding agent.
fx.sh1 month ago316 ptsView detailsJoin discussion
Field measurements of neighborhood-scale air temperature impacts of data centers
asmedigitalcollection.asme.org1 month ago310 ptsView detailsJoin discussion
Claude Code May–August 2026 weekly limits promotion
support.claude.com1 month ago295 ptsView detailsJoin discussion
Babies born under sugar rationing grew into adults with lower cancer risk
theconversation.com1 month ago292 ptsView detailsJoin discussion
Palomar: A registry of Lean verified mathematics
terrytao.wordpress.com1 month ago185 ptsView detailsJoin discussion
AI usage patterns in software teams
linear.app1 month ago187 ptsView detailsJoin discussion
Show HN: Automatically detect and patch walking-dead states in Sierra games
github.com1 month ago158 ptsView detailsJoin discussion
GLM-5.3 Artificial Analysis Benchmarks
artificialanalysis.ai1 month ago152 ptsView detailsJoin discussion
DFlash 2: Keep Drafting Parallel
inco.ai1 month ago98 ptsView detailsJoin discussion
Launch HN: machine0 (YC S26) – Persistent CPU and GPU VMs from the CLI
machine0.io1 month ago83 ptsView detailsJoin discussion
Companies promote incompetent employees to management to limit damage they do
lawsofsoftwareengineering.com1 month ago79 ptsView detailsJoin discussion
I used to be excited about new tech, but I rarely am anymore
82mhz.net1 month ago72 ptsView detailsJoin discussion
Show HN: Openleetcode – Local LeetCode runner where tests live in the repo
github.com1 month ago71 ptsView detailsJoin discussion
Show HN: Shoehorn – Quantize any model down to run on your machine
notactuallytreyanastasio.github.io1 month ago46 ptsView detailsJoin discussion
The National Park Service Is Using Flock. Rangers Are Pissed
404media.co1 month ago45 ptsView detailsJoin discussion
AI won't solve the work-theater problem
think-twice.me1 month ago42 ptsView detailsJoin discussion
nytimes.com1 month ago39 ptsView detailsJoin discussion
Micron, SK Commit Billions to RAM Capacity, but Almost Nothing Lands Before 2028
storagereview.com1 month ago35 ptsView detailsJoin discussion
ChatGPT has almost stopped citing Reddit
promptwatch.com1 month ago32 ptsView detailsJoin discussion
A group of Gandalfs protest outside the home of Peter Thiel in Argentina
dangerousminds.net1 month ago29 ptsView detailsJoin discussion
Show HN: macOS data protection keychain for Electron apps
github.com1 month ago24 ptsView detailsJoin discussion
Show HN: I canceled my AI code reviewer and wrote a free local one
github.com1 month ago22 ptsView detailsJoin discussion
Democracy vs. the machine: birth of digital age,the warnings that were ignored
theguardian.com1 month ago22 ptsView detailsJoin discussion
Palantir Leads AI Data Deal with USA Today Sparking a Newsroom Revolt
forbes.com1 month ago21 ptsView detailsJoin discussion
The Microsoft Rebrand Registry
msrebrandregistry.com1 month ago19 ptsView detailsJoin discussion
AI Alignment as a Thought-Terminating Cliche
borretti.me1 month ago19 ptsView detailsJoin discussion
devin.ai1 month ago18 ptsView detailsJoin discussion
200B Tokens Later: A Month of Letting AI Agents Decompile MW2
momo5502.com1 month ago18 ptsView detailsJoin discussion
s-1.vercel.app1 month ago18 ptsView detailsJoin discussion
What We Learned Moving Our Agent Loops from Anthropic to GLM
getunblocked.com1 month ago18 ptsView detailsJoin discussion
Elon Musk made flying worse so Palantir could profit
On August 6th, the Minneapolis Air Route Traffic Control Center lost radar and communications for around two hours. The outage disrupted more than 1,100 flights across the center's 330,000 square mile, nine-state airspace sector. Two days earlier, on August 4th, President Donald Trump departed the White House inside h…
theverge.com1 month ago17 ptsView detailsJoin discussion
Qwen3.8-27B make medium the default effort level instead of xhigh
github.com1 month ago16 ptsView detailsJoin discussion
Muse Glimmer is a memory hierarchy disguised as a 30B Transformer
abstractextraordinary.com1 month ago16 ptsView detailsJoin discussion
OpenAI Is Slowing Down Its AI Training
time.com1 month ago14 ptsView detailsJoin discussion
The most influential economist is oddly unconvincing
economist.com1 month ago13 ptsView detailsJoin discussion
U.S. to tell partners they must pick sides in AI race with China
cnbc.com1 month ago13 ptsView detailsJoin discussion
Show HN: PantheonGPU – GPU health testing and AI workload benchmarking
pantheongpu.com1 month ago13 ptsView detailsJoin discussion
Association of Spicy Chilli Food Consumption with All-Cause Mortality
pubmed.ncbi.nlm.nih.gov1 month ago12 ptsView detailsJoin discussion
AI Is Upending One of Finance's Cushiest Jobs
bloomberg.com1 month ago11 ptsView detailsJoin discussion
The market for used EVs 'is so hot'
grist.org1 month ago11 ptsView detailsJoin discussion
State Farm defense lawyers admit AI generated fake cases in LA lawsuit
calmatters.org1 month ago11 ptsView detailsJoin discussion
Pilots, Flight Attendants Top in the List of Radiation-Related Cancer Deaths
hms.harvard.edu1 month ago11 ptsView detailsJoin discussion
Mythic's analog compute-in-memory architecture
mythic.ai1 month ago10 ptsView detailsJoin discussion
NASA's IXPE detects 80% X-ray polarization from magnetar 1E 1547.0-5408
nature.com1 month ago10 ptsView detailsJoin discussion
AI to help planes avoid climate-warming 'sky graffiti'
bbc.com1 month ago10 ptsView detailsJoin discussion
- Primary source
ChatGPT Ads expands across Europe
ChatGPT Ads is expanding to 31 European markets. Learn how advertisers can reach people as they explore, compare options, and make decisions.
openai.com1 month agoView details
- Primary source
Strengthening democratic oversight in national security
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expertise.
openai.com1 month agoView details
- Primary source
Introducing ChatGPT for Teens: Built for learning, backed by protections
ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents.
openai.com1 month agoView details
- Primary source
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
openai.com1 month agoView details
- Primary source
Partnering with CodeAI to prepare the first AI generation
OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
openai.com1 month agoView details
- Primary source
Asana cleared 5 years of engineering work in 2 weeks with Codex
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
openai.com1 month agoView details
NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an Apache-2.0 project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands, with no intermediate ONNX export. The build emits a versioned .bundle artifact that runs through native C++ task APIs, so…
marktechpost.com1 month agoView details
Meet SAM (Sovereign Agent Mesh): A Zero-Config, Zero-Trust P2P Network for AI Agents
Google has open-sourced SAM (Sovereign Agent Mesh) under Apache-2.0 — and it has nothing to do with Segment Anything. SAM is a zero-config, zero-trust P2P overlay that lets autonomous agents discover and call each other's MCP tools across cloud, on-prem, laptop and edge environments, without exposing a single internal…
marktechpost.com1 month agoView details
Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to…
marktechpost.com1 month agoView details
Robin Williams’ Instagram account brought back to fight ‘AI abuse’
Robin Williams' children are taking over their father's Instagram account after his daughter spoke out against the use of his AI likeness, as reported earlier by The Wrap. In a post on Tuesday, Zak, Zelda, and Cody Williams write that they want the late actor's Instagram profile to be a "safe, trusted place where the…
theverge.com1 month agoView details
OpenAI lays out new security changes after its AI hacked Hugging Face
OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks co…
theverge.com1 month agoView details
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.
wired.com1 month agoView details
Firefox’s Smart Window promises a better AI browser
Starting today, AI chats in Firefox's Smart Window AI browsing mode can pull from current web info and show source links in chat responses through a partnership with Exa. Smart Window can also now automatically suggest tab groups and show visual previews of pages you previously visited when you search your browsing hi…
theverge.com1 month agoView details
Google’s Pet Memory forgot who my cats are
One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, I thought th…
theverge.com1 month agoView details
ChatGPT is getting a dedicated mode for teens
ChatGPT for Teens includes safeguards and parental controls. | Image: OpenAI OpenAI is introducing a dedicated ChatGPT mode for teenagers, combining existing youth safeguards and new safety features under one roof. The launch comes amid mounting public scrutiny over how AI tools affect younger users, as other platform…
theverge.com1 month agoView details
Apple’s camera-equipped AirPods appear in leaked video
The AirPods in the video look like a chunkier version of Apple’s AirPods Pro 3. | Image: Apple / MacRumors We may have our first glimpse of Apple's rumored camera-equipped AirPods, thanks to a video that MacRumors found in the macOS Tahoe 26.7 Release Candidate. The short video clip features a man - who is wearing the…
theverge.com1 month agoView details
The Powerful Chinese AI Model Experts Warned About Is Here
Z.ai’s latest AI model release could help companies secure their systems—or find its way into the hands of hackers.
wired.com1 month agoView details
J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers
arXiv:2608.17063v1 Announce Type: cross Abstract: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the mod…
arxiv.org1 month agoView details
When AI Designs AI: Innovation or Imitation?
arXiv:2608.17471v1 Announce Type: new Abstract: Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic desi…
arxiv.org1 month agoView details
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
arXiv:2608.17530v1 Announce Type: new Abstract: Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate c…
arxiv.org1 month agoView details
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation
arXiv:2608.17588v1 Announce Type: new Abstract: Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate sol…
arxiv.org1 month agoView details
arXiv:2608.17911v1 Announce Type: new Abstract: As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability problem in long-term conversational memo…
arxiv.org1 month agoView details
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
arXiv:2608.17330v1 Announce Type: new Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, mul…
arxiv.org1 month agoView details
arXiv:2608.17684v1 Announce Type: new Abstract: Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether learned behavior preserves previously correct behavior or security. We audit SkillOpt, Agent Workflow Memory (AWM), and ReasoningBan…
arxiv.org1 month agoView details
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
arXiv:2608.17468v1 Announce Type: new Abstract: Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying direct…
arxiv.org1 month agoView details
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
arXiv:2608.17856v1 Announce Type: new Abstract: Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Mo…
arxiv.org1 month agoView details
Children, but not language models, show accelerating returns in word learning
arXiv:2608.17120v1 Announce Type: new Abstract: Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerati…
arxiv.org1 month agoView details
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
arXiv:2608.17336v1 Announce Type: new Abstract: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial…
arxiv.org1 month agoView details
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-ag…
arxiv.org1 month agoView details
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
arXiv:2608.17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defen…
arxiv.org1 month agoView details
Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context
arXiv:2608.17499v1 Announce Type: new Abstract: User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repai…
arxiv.org1 month agoView details
The politics of postmortem privacy
arXiv:2608.16905v1 Announce Type: cross Abstract: While the existence of postmortem privacy is increasingly acknowledged (such as the protection of the presence of deceased within digital spaces), far less attention has been paid to its internal instability: its scope (the extent of its application), justificatory fou…
arxiv.org1 month agoView details
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
arXiv:2608.17749v1 Announce Type: new Abstract: Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat…
arxiv.org1 month agoView details
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
arXiv:2608.17634v1 Announce Type: new Abstract: The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only th…
arxiv.org1 month agoView details
arXiv:2608.17168v1 Announce Type: new Abstract: Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of leg…
arxiv.org1 month agoView details
Evaluating the Diversity of AI-Generated Content with Diversity Profiles
arXiv:2608.17731v1 Announce Type: new Abstract: Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarit…
arxiv.org1 month agoView details
ArborMem: Navigating Interaction States with Memory Forests
arXiv:2608.17534v1 Announce Type: new Abstract: Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions. Existing methods improve access to conversational history through long-context processing,…
arxiv.org1 month agoView details
Token Optimization and Context Window Management in Multi-Agent AI Workflows
arXiv:2608.17188v1 Announce Type: new Abstract: Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents a practitioner framework for token optimization and context-window management, grounded in an internal production dashboard that ext…
arxiv.org1 month agoView details
An Investigation of the NeurIPS and ICML 2025 Position Tracks
arXiv:2608.16894v1 Announce Type: cross Abstract: ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper ar…
arxiv.org1 month agoView details
AutoResearch: Insight In, Hallucination Out
arXiv:2608.17906v1 Announce Type: new Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Id…
arxiv.org1 month agoView details
EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection
arXiv:2608.17933v1 Announce Type: new Abstract: Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend he…
arxiv.org1 month agoView details
Polaris: Learning to Generate Table Descriptions from Retrieval Feedback
arXiv:2608.17171v1 Announce Type: new Abstract: Many table-centric NLP tasks such as NL2SQL first retrieve relevant tables from large collections using keyword search. Recent work uses LLMs to generate natural-language table descriptions to improve retrieval, but they are typically optimized for fluency rather than re…
arxiv.org1 month agoView details
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
arXiv:2608.17638v1 Announce Type: new Abstract: What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own r…
arxiv.org1 month agoView details
LLM-Derived Preference Judgments Are Not Self-Consistent
arXiv:2608.17644v1 Announce Type: new Abstract: Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments…
arxiv.org1 month agoView details
Chain-of-Experience for Continual LLM Improvement
arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we re…
arxiv.org1 month agoView details
Accuracy and Robustness of Model Cascades Under Data Perturbations
arXiv:2608.17711v1 Announce Type: new Abstract: Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a la…
arxiv.org1 month agoView details
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
arXiv:2608.17800v1 Announce Type: new Abstract: Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that r…
arxiv.org1 month agoView details
arXiv:2608.17352v1 Announce Type: new Abstract: Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Gra…
arxiv.org1 month agoView details
arXiv:2608.17583v1 Announce Type: new Abstract: Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing…
arxiv.org1 month agoView details
Cross-Model Memory Transfer via Target-Side Reader Adaptation
arXiv:2608.17050v1 Announce Type: new Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric a…
arxiv.org1 month agoView details
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method
arXiv:2608.17536v1 Announce Type: new Abstract: Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in…
arxiv.org1 month agoView details
arXiv:2608.17950v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true…
arxiv.org1 month agoView details
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
arXiv:2608.17994v1 Announce Type: new Abstract: Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a…
arxiv.org1 month agoView details
Adaptive Policy Portfolios for Robust Markov Decision Processes
arXiv:2608.17929v1 Announce Type: new Abstract: Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memo…
arxiv.org1 month agoView details
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
arXiv:2608.17150v1 Announce Type: cross Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not…
arxiv.org1 month agoView details
AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction
arXiv:2608.17184v1 Announce Type: new Abstract: Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (L…
arxiv.org1 month agoView details
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
arXiv:2608.18011v1 Announce Type: new Abstract: Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competitio…
arxiv.org1 month agoView details
The Plot Thins: Uniformity and Linearity in Literary Summaries
arXiv:2608.17218v1 Announce Type: new Abstract: Works of literature are complicated; they balance plot, suspense, surprise, and artistic expression. Summaries of literature prioritize plot, and therefore may deviate from their sources. Using a combination of manual and LLM-based annotation, we construct a dataset mapp…
arxiv.org1 month agoView details
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
arXiv:2608.17665v1 Announce Type: new Abstract: LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or co…
arxiv.org1 month agoView details
Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents
arXiv:2608.17153v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. No…
arxiv.org1 month agoView details
Procedural Content Metageneration via Program Search and Continual Abstraction Discovery
arXiv:2608.17947v1 Announce Type: new Abstract: Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Pytho…
arxiv.org1 month agoView details
Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence
arXiv:2608.16975v1 Announce Type: new Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. This amb…
arxiv.org1 month agoView details
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
arXiv:2608.17756v1 Announce Type: new Abstract: Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluati…
arxiv.org1 month agoView details
The Problem Is the Problem: Towards Scalable Mathematical Discovery
arXiv:2608.16977v1 Announce Type: new Abstract: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore centra…
arxiv.org1 month agoView details
Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents
arXiv:2608.17718v1 Announce Type: new Abstract: Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized…
arxiv.org1 month agoView details
There is No Theoretical Curse of Multilinguality For Embedding Space Structure
arXiv:2608.17088v1 Announce Type: new Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model…
arxiv.org1 month agoView details
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
arXiv:2608.17223v1 Announce Type: new Abstract: Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a 49,799-article corpus across 16 feature-model…
arxiv.org1 month agoView details
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
arXiv:2608.17501v1 Announce Type: new Abstract: Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI…
arxiv.org1 month agoView details
arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate u…
arxiv.org1 month agoView details
Which Source Wins? Task-Dependent Reliance in Vision-Language Models
arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the te…
arxiv.org1 month agoView details
Agent Lightning v1.0: Towards Harnessed Agentic RL
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through a…
arxiv.org1 month agoView details
arXiv:2608.17051v1 Announce Type: new Abstract: Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We a…
arxiv.org1 month agoView details
Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention
arXiv:2608.17288v1 Announce Type: new Abstract: GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model whether strong token features should reinforce or suppress one another. We introduce Q-Interference, a f…
arxiv.org1 month agoView details
MoNe: Modular Neural Memory for Efficient Long Context Inference
arXiv:2608.17616v1 Announce Type: new Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-…
arxiv.org1 month agoView details
arXiv:2608.17781v1 Announce Type: new Abstract: ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can…
arxiv.org1 month agoView details
SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
arXiv:2608.17931v1 Announce Type: new Abstract: Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world app…
arxiv.org1 month agoView details
Grading Needs a Rubric, Not Intelligence
arXiv:2608.17938v1 Announce Type: new Abstract: Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at…
arxiv.org1 month agoView details
arXiv:2608.18041v1 Announce Type: new Abstract: Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention weights refine that count, which sums every writer in the corpus together. This paper claims a second parameter, phase…
arxiv.org1 month agoView details
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
arXiv:2608.17443v1 Announce Type: new Abstract: Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs)…
arxiv.org1 month agoView details
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools
arXiv:2608.17007v1 Announce Type: new Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire in…
arxiv.org1 month agoView details
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
arXiv:2608.17959v1 Announce Type: new Abstract: State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-depend…
arxiv.org1 month agoView details
arXiv:2608.17625v1 Announce Type: new Abstract: Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms. We deliv…
arxiv.org1 month agoView details
arXiv:2608.17124v1 Announce Type: new Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so th…
arxiv.org1 month agoView details
Toward Personal Intelligence Through Cooperative Observation
arXiv:2608.17128v1 Announce Type: new Abstract: A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a boun…
arxiv.org1 month agoView details
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents
arXiv:2608.16890v1 Announce Type: new Abstract: Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontie…
arxiv.org1 month agoView details
arXiv:2608.16891v1 Announce Type: new Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it…
arxiv.org1 month agoView details
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
arXiv:2608.17075v1 Announce Type: new Abstract: Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correc…
arxiv.org1 month agoView details
arXiv:2608.17096v1 Announce Type: new Abstract: The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with…
arxiv.org1 month agoView details
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
arXiv:2608.17170v1 Announce Type: new Abstract: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an autom…
arxiv.org1 month agoView details
LLM-Only PDDL Domain Repair with Open-Weight Models
arXiv:2608.17341v1 Announce Type: new Abstract: AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such model…
arxiv.org1 month agoView details
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
arXiv:2608.17741v1 Announce Type: new Abstract: OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entai…
arxiv.org1 month agoView details
Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
arXiv:2608.17687v1 Announce Type: new Abstract: Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is…
arxiv.org1 month agoView details
FedPref: Federated Preference Learning for Structured Radiology Report Extraction
arXiv:2608.16971v1 Announce Type: new Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evi…
arxiv.org1 month agoView details
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
arXiv:2608.17102v1 Announce Type: new Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared af…
arxiv.org1 month agoView details
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large…
arxiv.org1 month agoView details
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
arXiv:2608.17247v1 Announce Type: new Abstract: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, the…
arxiv.org1 month agoView details
arXiv:2608.17270v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similar…
arxiv.org1 month agoView details
arXiv:2608.17574v1 Announce Type: new Abstract: How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein…
arxiv.org1 month agoView details
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
arXiv:2608.17744v1 Announce Type: new Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only…
arxiv.org1 month agoView details
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
arXiv:2608.16956v1 Announce Type: cross Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered…
arxiv.org1 month agoView details
ASI-Bench: At the Dawn of Artificial Superintelligence
arXiv:2608.17271v1 Announce Type: new Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on lear…
arxiv.org1 month agoView details
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
arXiv:2608.17282v1 Announce Type: new Abstract: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework tha…
arxiv.org1 month agoView details
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories d…
arxiv.org1 month agoView details
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
arXiv:2608.17299v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks pr…
arxiv.org1 month agoView details
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
arXiv:2608.17301v1 Announce Type: new Abstract: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems rema…
arxiv.org1 month agoView details
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents
arXiv:2608.17319v1 Announce Type: new Abstract: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignme…
arxiv.org1 month agoView details
ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation
arXiv:2608.17356v1 Announce Type: new Abstract: Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three de…
arxiv.org1 month agoView details
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
arXiv:2608.17379v1 Announce Type: new Abstract: We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over f…
arxiv.org1 month agoView details
An Investigation of Translationese in the Generations of Multilingual Large Language Models
arXiv:2608.17399v1 Announce Type: new Abstract: Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, i…
arxiv.org1 month agoView details
From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis
arXiv:2608.17454v1 Announce Type: new Abstract: This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level cov…
arxiv.org1 month agoView details
Effects of Answer Format Variation on Gender Bias in Large Language Models
arXiv:2608.17516v1 Announce Type: new Abstract: Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a…
arxiv.org1 month agoView details
arXiv:2608.17587v1 Announce Type: new Abstract: Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inferen…
arxiv.org1 month agoView details
arXiv:2608.17605v1 Announce Type: new Abstract: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context acro…
arxiv.org1 month agoView details
TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification
arXiv:2608.17795v1 Announce Type: new Abstract: Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database c…
arxiv.org1 month agoView details
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
arXiv:2608.17810v1 Announce Type: new Abstract: The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM perfor…
arxiv.org1 month agoView details
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this pape…
arxiv.org1 month agoView details
arXiv:2608.17843v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a con…
arxiv.org1 month agoView details
BayesPrompt: human readable prompts that make sense
arXiv:2608.17866v1 Announce Type: new Abstract: Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the perplexity of a given answer, however, consistently yield so-called pseudoprompts, unintelligible stri…
arxiv.org1 month agoView details
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
arXiv:2608.17895v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external d…
arxiv.org1 month agoView details
TokEval: A Tokenizer Evaluation Suite
arXiv:2608.18062v1 Announce Type: new Abstract: Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstr…
arxiv.org1 month agoView details
arXiv:2608.18072v1 Announce Type: new Abstract: Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictate…
arxiv.org1 month agoView details
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis
arXiv:2311.06273v1 Announce Type: cross Abstract: The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Recognizing its value across different areas, our study investigates ChatGPT's capacity to predict stock market movements using only…
arxiv.org1 month agoView details
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
arXiv:2602.14784v1 Announce Type: cross Abstract: Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return…
arxiv.org1 month agoView details
arXiv:2608.15382v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-c…
arxiv.org1 month agoView details
arXiv:2608.16909v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok…
arxiv.org1 month agoView details
SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback
arXiv:2608.16934v1 Announce Type: cross Abstract: RTL code generation is a critical stage in hardware design, and the emergence of agentic systems offers new opportunities to automate this process. To generate correct RTL code, agents must understand sequential behavior, including how signals evolve and propagate over…
arxiv.org1 month agoView details
Memory Is Communication: The Frontier Between Remembering and Signaling
arXiv:2608.17053v1 Announce Type: cross Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an a…
arxiv.org1 month agoView details
Uncertainty-Aware Decision Making in Multimodal Large Language Models
arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input qualit…
arxiv.org1 month agoView details
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
arXiv:2608.17325v1 Announce Type: new Abstract: Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across languages. In this work, we analyze what tokens are learned when tokenization is jointly optim…
arxiv.org1 month agoView details
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
arXiv:2608.17809v1 Announce Type: new Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLM…
arxiv.org1 month agoView details
arXiv:2608.17979v1 Announce Type: new Abstract: Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI-assisted writ…
arxiv.org1 month agoView details
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
arXiv:2608.17393v1 Announce Type: new Abstract: Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradi…
arxiv.org1 month agoView details
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations
arXiv:2608.17433v1 Announce Type: new Abstract: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the sam…
arxiv.org1 month agoView details
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression
arXiv:2608.17434v1 Announce Type: new Abstract: We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variation budget A, and output bound B. For this O(L w^2)-parameterized architecture, the known lower and upper bounds differ by one factor…
arxiv.org1 month agoView details
Models
View all models →orcarouter/Qwen3.8-27B-Uncensored-GGUF
image-text-to-text · gguf · abliterated · qwen
huggingface.co1 month ago310 ptsView details
peculiar-ragdoll/Qwen-Sharp-Chat-Templates
mlx · jinja · chat-template
huggingface.co1 month ago280 ptsView details
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
image-text-to-text · gguf · qwen3_5 · unsloth
huggingface.co1 month ago272 ptsView details
text-generation · transformers · safetensors · qwen3
huggingface.co1 month ago241 ptsView details
text-generation · transformers · safetensors · qwen3
huggingface.co1 month ago181 ptsView details
incoai/Qwen3.8-27B-DFlash2-GGUF
text-generation · llama.cpp · gguf · dflash2
huggingface.co1 month ago121 ptsView details
z-lab/Qwen3.8-27B-DFlash2-GGUF
text-generation · llama.cpp · gguf · dflash2
huggingface.co1 month ago85 ptsView details
esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
text-generation · gguf · nvfp4 · qwen3.8
huggingface.co1 month ago54 ptsView details
mlasli/Qwen3.8-27B-Heretic-Uncensored-Q4_K_M-GGUF
text-generation · gguf · qwen3.8 · qwen3
huggingface.co1 month ago2 ptsView details
text-generation · transformers · safetensors · gpt2
huggingface.co1 month ago1 ptsView details
text-generation · transformers · safetensors · qwen3
huggingface.co1 month ago1 ptsView details
text-generation · transformers · safetensors · llama
huggingface.co1 month agoView details
MohamedAhmedAE/llava-medical-8B-clip-vit_kaggle-stage2
safetensors · llava · region:us
huggingface.co1 month agoView details
text-generation · transformers · safetensors · llama
huggingface.co1 month agoView details
Open source
View all open source →## [3.3.0](https://github.com/openai/openai-python/compare/v3.2.0...v3.3.0) (2026-08-18) ### Features * support named data-residency endpoints ([#3646](https://github.com/openai/openai-python/issues/3646)) ([11ee914](https://github.com/openai/openai-python/commit/11ee91475694d9c…
github.com1 month agoView details
langchain-ai/langchain langchain-openai==1.5.2
Changes since langchain-openai==1.5.1 release(openai): 1.5.2 (#39719) fix(openai): preserve reasoning item boundaries (#39278) release(openai): 1.5.2a1 (#39709) feat(openai): extract gateway metadata from response headers when available (#39706) chore(openai): update snapshots (…
github.com1 month agoView details
modelcontextprotocol/servers 2026.8.18
# Release : v2026.8.18 ## Updated packages - @modelcontextprotocol/server-everything@2026.8.18 - mcp-server-time@2026.8.18 - mcp-server-fetch@2026.8.18 - mcp-server-git@2026.8.18
github.com1 month agoView details
<details open> ci : Update OpenVINO to 2026.3, skip nemotron-h rollback test (#27292) * update to ov-2026.3, update device drivers * ci: skip nemotron-h rollback test on OpenVINO The OpenVINO backend does not support SSM_SCAN, so the Nemotron-H recurrent state rollback graph is…
github.com1 month agoView details
<details open> mtmd: fix LFM2 image tiling threshold (#27057) * mtmd: fix LFM2 image tiling threshold * refactor testing * fix * fix on windows --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Ap…
github.com1 month agoView details
> [!NOTE] > Semantic versioning is still work in progress. > More info can be found in https://github.com/ggml-org/ggml/discussions/1579 **Nightly build:** [b10485](https://github.com/ggml-org/llama.cpp/releases/tag/b10485) ## Change log since v0.1.1 1511ce3bc sync : ggml da786d…
github.com1 month agoView details