Archive / 2026-08-11
August 11, 2026
News
View all news →France to ban unsolicited telemarketing calls
lemonde.fr1 month ago1062 ptsView detailsJoin discussion
Stealing Reasoning Traces from Proprietary LLM APIs
stolen-thoughts.com1 month ago697 ptsView detailsJoin discussion
OpenAI’s head of ethics leaves less than a year after joining
ft.com1 month ago511 ptsView detailsJoin discussion
Go is an ideal language for AI-assisted software engineering
developers.googleblog.com1 month ago430 ptsView detailsJoin discussion
llama.app1 month ago364 ptsView detailsJoin discussion
x.ai1 month ago338 ptsView detailsJoin discussion
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
github.com1 month ago307 ptsView detailsJoin discussion
WorldClaw Agentic 3D open-world generation at scale
tencent-hunyuan.github.io1 month ago277 ptsView detailsJoin discussion
US hires over 2k video gamers as air traffic controllers
cbsnews.com1 month ago200 ptsView detailsJoin discussion
Company Offering '100% Human-Written, Never AI' Medical Research Is 100% AI
404media.co1 month ago195 ptsView detailsJoin discussion
The lifesaving secret hidden inside a horseshoe crab's blue blood
whdh.com1 month ago103 ptsView detailsJoin discussion
Lean Eval for Alignment on Faithfulness
millenniumresearch.ai1 month ago103 ptsView detailsJoin discussion
$580M undersea cable rerouted to avoid the grave of Dobby the House Elf
tomshardware.com1 month ago102 ptsView detailsJoin discussion
Why Did OpenAI's Head of Ethics Chloé Bakalar Leave?
aimagazine.com1 month ago87 ptsView detailsJoin discussion
Why your Amazon order confirmation emails have become so unhelpful
Earlier this summer, Amazon customers began noticing that emails related to their online orders looked sparse: Order confirmation emails didn't name specific items anymore, and instead listed only item categories. "Your Beauty item is confirmed!" an email about my retainer cleaning tablets read. Shoppers have posted o…
theverge.com1 month ago78 ptsView detailsJoin discussion
DARPA heavy lift challenge ends with winner at a 3.84:1 payload to weight ratio
dronexl.co1 month ago71 ptsView detailsJoin discussion
Half of Europe's towns and villages have fewer residents than 60 years ago
correctiv.org1 month ago64 ptsView detailsJoin discussion
x.ai1 month ago59 ptsView detailsJoin discussion
Zuckerberg's superyacht ignored emergency channel, failed to aid stranded boat
arstechnica.com1 month ago59 ptsView detailsJoin discussion
Suzanne: AI tool for designing and manufacturing physical products
suzanne3d.com1 month ago51 ptsView detailsJoin discussion
doi.org1 month ago48 ptsView detailsJoin discussion
Launch HN: Keet (YC S24) – An app to create video courses on anything
trykeet.com1 month ago47 ptsView detailsJoin discussion
Closing Canario Terminal source code
rapha.land1 month ago46 ptsView detailsJoin discussion
Show HN: Lambdock – Wayland-native GTK4 dock with a live Lisp REPL
codeberg.org1 month ago44 ptsView detailsJoin discussion
Show HN: OJCP – an open protocol for agent-consumable job data
ojcp.dev1 month ago39 ptsView detailsJoin discussion
Claude Code is leaking real email address as a User-Agent string in curl command
github.com1 month ago38 ptsView detailsJoin discussion
The whole of PyTorch on one page
tensor.khalilli.ai1 month ago35 ptsView detailsJoin discussion
Gemini becomes Google's fastest-growing product ever as it hits 1B users
arstechnica.com1 month ago29 ptsView detailsJoin discussion
FDA report reveals where Taylor Farms sent its lettuce–and it raises questions
arstechnica.com1 month ago28 ptsView detailsJoin discussion
What happens when the AI bubble pops?
thehustle.co1 month ago24 ptsView detailsJoin discussion
AI Is Solving CTF Challenges in Minutes
simulationslabs.com1 month ago21 ptsView detailsJoin discussion
Statin use vs. death rate from cardiovascular diseases (2019)
ourworldindata.org1 month ago20 ptsView detailsJoin discussion
Show HN: TermDOM – HTML, CSS and JavaScript (With a Real DOM) for TUIs and CLIs
github.com1 month ago17 ptsView detailsJoin discussion
New Orleans is using AI to triage 911 calls in case of backlog
consumerrights.wiki1 month ago16 ptsView detailsJoin discussion
4Chan Chuds Used AI to Clothe Her. She Fought Back (2024)
rollingstone.com1 month ago15 ptsView detailsJoin discussion
unsloth.ai1 month ago14 ptsView detailsJoin discussion
Linus Torvalds says AI has made 'huge' Linux kernel updates the new normal
theregister.com1 month ago13 ptsView detailsJoin discussion
Crew, a multiplayer workspace for humans and AI agents to work together
github.com1 month ago12 ptsView detailsJoin discussion
Review of Studies on the Level of Cannabis Use and Risk of Psychosis
pmc.ncbi.nlm.nih.gov1 month ago12 ptsView detailsJoin discussion
greggman.github.io1 month ago12 ptsView detailsJoin discussion
FlightAware Sues Kalshi over Flight Cancellation Prediction Markets
wsj.com1 month ago11 ptsView detailsJoin discussion
blog.oxplot.com1 month ago11 ptsView detailsJoin discussion
I backtested my own stock rankings. They lost to the index
holderdashboard.com1 month ago10 ptsView detailsJoin discussion
- Primary source
How RingCentral builds AI-native work from engineering to ops
See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.
openai.com1 month agoView details
- Primary source
OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
openai.com1 month agoView details
- Primary source
Daybreak models are now available on AWS
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
openai.com1 month agoView details
- Primary source
Advancing AMIE towards expert-level audio-visual clinical consultations
Health & Bioscience
research.google1 month agoView details
LTX-2.5 brings frontier video generation to local NVIDIA hardware: 6.8-second clips, native multishot, day-one ComfyUI, open weights. The post The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model appeared first on MarkTechPost.
marktechpost.com1 month agoView details
In this tutorial, we build a complete quantitative backtesting workflow with OctoBot and OctoBot-Script while keeping the environment isolated from Colab’s preinstalled dependencies. We configure a rule-based trading strategy that combines RSI-based oversold signals, EMA trend confirmation, and ATR-driven adaptive sto…
marktechpost.com1 month agoView details
webAI has released TwIL-LM, a family of formal-logic models at 1.7B and 3B parameters that translate English into first-order logic and check whether conclusions follow from premises. The 3B runs on CPU or 4GB of VRAM; the 1.7B downloads at 1.06GB. Both ship under a non-commercial license. The model card also shows th…
marktechpost.com1 month agoView details
Implementing a MiniMax-H3 Multimodal Video and Audio Generation Pipeline with ComfyUI APIs
In this comprehensive guide, we demonstrate how to implement a complete, programmable MiniMax-H3 multimodal generation pipeline. By leveraging ComfyUI as a headless backend, we walk through setting up an automated inference environment that handles hardware profiling, model weight downloading, dynamic graph constructi…
marktechpost.com1 month agoView details
Saber denies replacing Rideshare Stimulator’s writers with ChatGPT
After a former lead writer claimed Saber "replaced me with ChatGPT," CEO Matthew Karch now claims, "Neither Saber nor Unigine have replaced any writers with AI," for the Rideshare "Stimulator" game announced last month, developed by Unigine. The writer, Stella Sacco, says differently, however, posting on Bluesky that…
theverge.com1 month agoView details
ChatGPT and Gemini both just passed 1 billion users
That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milest…
theverge.com1 month agoView details
Another OpenAI executive takes off
Brad Lightcap, OpenAI's special projects lead and the company's former COO, announced his departure after an eight-year stint at the AI lab. In an internal memo he later posted to X, Lightcap told colleagues he'd be starting "something new." "Over the last few months, I've been focused on the next horizon and what wou…
theverge.com1 month agoView details
Made by Google 2026: all the Pixel news and announcements
On August 12, 2026, Google revealed a bunch of new Pixel devices. The colorful Pixel 11 lineup comes with upgraded cameras and performance, with the Pro models offering a built-in LED ring that lights up for Google’s Gemini AI and other features. Google also showed off its next-gen Pixel Fold featuring thinner bezels,…
theverge.com1 month agoView details
Apple could help you prove your iPhone photos aren’t deepfakes
Apple is seemingly developing an iOS feature that can verify when a photograph was taken using an iPhone camera. 9to5Mac reports that the iOS 27 beta 5 includes code references for an "Apple Reference Image" system that can embed provenance metadata into iPhone photographs at the point of capture - enabling users to p…
theverge.com1 month agoView details
‘Zoomsday’ hack uncovered using fewer than 20 AI prompts
Zoom has patched a major security vulnerability that could allow an attacker to hijack anyone's device during a meeting. In a blog post on Tuesday, researchers at A Security say they uncovered the flaw using "fewer than 20 prompts on publicly available AI models," as reported earlier by Wired. The exploit involved Zoo…
theverge.com1 month agoView details
Spotify says it won’t recommend music from ‘AI Personas’
Spotify will soon label AI artists and remove their music from your recommendations. The change, which will start rolling out in mid-September, means you'll see an "AI Persona" badge on an artist's profile across the app if they do "not represent a real person." The music streaming platform will allow artists to discl…
theverge.com1 month agoView details
A Zoom Screen-Sharing Bug Let Anyone Take Over Other Devices on a Call
Researchers say it took fewer than 20 prompts for a public AI tool to find a flaw (now fixed) allowing anyone on a Zoom call to hijack another participants’ device.
wired.com1 month agoView details
Claude will apply invisible watermarks to AI text and images
Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a…
theverge.com1 month agoView details
The AI takeover of mathematics has begun
Mathematician James Maynard has spent a lot of time this past year "soul searching." A professor at the University of Oxford and winner of the prestigious Fields Medal, Maynard told The Verge he's been grappling with the future of his field as the traditionally slow-moving discipline hurries to adapt to AI. Days befor…
theverge.com1 month agoView details
A New Trick Reveals AI Models’ Inner Thoughts
Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.
wired.com1 month agoView details
AI Is Dead. Organoids Are Alive
Mini human brains are being grown in labs all over the world. Soon, they could outthink neural networks.
wired.com1 month agoView details
AI Is Helping Solve the Intricate Genetic Puzzle of Schizophrenia
Recent findings provide one of the most detailed pictures to date of the genetic architecture of schizophrenia, opening up new avenues for research into the disorder.
wired.com1 month agoView details
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We intr…
arxiv.org1 month agoView details
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat RAG diaries and Markdown logs optimize needle retrieval but under-serve currency, provenance, and ordered temporal rea…
arxiv.org1 month agoView details
Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
arXiv:2608.10315v1 Announce Type: new Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a…
arxiv.org1 month agoView details
Towards an Argumentative Foundation for Evaluative AI
arXiv:2608.07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) arg…
arxiv.org1 month agoView details
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a curriculum-framework paper, grounded in a structured narrative review of labor-market and software-engineering evidence and illustrate…
arxiv.org1 month agoView details
arXiv:2608.10939v1 Announce Type: new Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference po…
arxiv.org1 month agoView details
arXiv:2608.11044v1 Announce Type: new Abstract: Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengt…
arxiv.org1 month agoView details
Towards Researcher Agents for Knowledge-Graph Question Answering
arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and seman…
arxiv.org1 month agoView details
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
arXiv:2608.07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking rich comparative s…
arxiv.org1 month agoView details
Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
arXiv:2608.07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executions are non-deterministic, and AI workloads bring high-dimensional inputs and non-linear operations that defeat fuzzing and symbolic…
arxiv.org1 month agoView details
arXiv:2608.07476v1 Announce Type: new Abstract: We develop a formal framework for constructing canonical interpretations from plural structure theories. A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose admissible interpretation family collects all gl…
arxiv.org1 month agoView details
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
arXiv:2608.10678v1 Announce Type: new Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3)…
arxiv.org1 month agoView details
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion…
arxiv.org1 month agoView details
Legal Responsibilities Using Autonomous Agents For Artificial Intelligence
arXiv:2608.08022v1 Announce Type: new Abstract: Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant criminal or negligent d…
arxiv.org1 month agoView details
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The worl…
arxiv.org1 month agoView details
Emotion in an active inference model of human driving
arXiv:2608.07480v1 Announce Type: new Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It has been successfully applied across biological and artificial systems, including recent work on human driving. However,…
arxiv.org1 month agoView details
Training Variable Long Sequences with Data-Centric Parallel
arXiv:2608.07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches use static configurations that cause workload imbalance low efficiency,…
arxiv.org1 month agoView details
arXiv:2608.07542v1 Announce Type: new Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of whichever metric they optimise rather than testing the hypotheses that motivate the experiments. We address this structur…
arxiv.org1 month agoView details
arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasonin…
arxiv.org1 month agoView details
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed char…
arxiv.org1 month agoView details
How Robust Are LLMs to Vietnamese Dialects?
arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through d…
arxiv.org1 month agoView details
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show thi…
arxiv.org1 month agoView details
arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or position-dependent rotations into Transformer representa…
arxiv.org1 month agoView details
Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue
arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape tru…
arxiv.org1 month agoView details
Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems
arXiv:2608.10216v1 Announce Type: new Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question: "Does this text still mean the same thin…
arxiv.org1 month agoView details
arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $\beta$-fraction of shifted target traffic with at most an $\alpha$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Fl…
arxiv.org1 month agoView details
arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventi…
arxiv.org1 month agoView details
Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases
arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities,…
arxiv.org1 month agoView details
Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So
arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. We show it is functional. A 12-layer model computes in two phases. Throu…
arxiv.org1 month agoView details
Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation
arXiv:2608.07949v1 Announce Type: new Abstract: Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneous sources, but prov…
arxiv.org1 month agoView details
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems should be built. We at…
arxiv.org1 month agoView details
Simplex Relaxation for Discrete Diffusion
arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask whether its training objective and reverse tr…
arxiv.org1 month agoView details
AndroidReality: How Far Are Mobile Agents from the Real World?
arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions. In this work, we introduce AndroidReal…
arxiv.org1 month agoView details
arXiv:2608.07627v1 Announce Type: new Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge…
arxiv.org1 month agoView details
Most biomedical publications show signs of LLM-assisted writing
arXiv:2608.10715v1 Announce Type: new Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decis…
arxiv.org1 month agoView details
Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems
arXiv:2608.07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and error propagation in…
arxiv.org1 month agoView details
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metr…
arxiv.org1 month agoView details
QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing
arXiv:2608.07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible complexity scope. We pre…
arxiv.org1 month agoView details
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply. The shape recurs: an origin's request-rate cap split across its edge locations, a licensed throughput cap ac…
arxiv.org1 month agoView details
MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales
arXiv:2608.10974v1 Announce Type: new Abstract: Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often along with reasoning on why those methods were chosen. We introduce MUSE (Mining Underlying Scientific Explanations),…
arxiv.org1 month agoView details
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic systems, a multi-component form of self…
arxiv.org1 month agoView details
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes…
arxiv.org1 month agoView details
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns
arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by fixed rules. Here, we present Agent-MD,…
arxiv.org1 month agoView details
Contextual Value Alignment via Multilayer Combinatorial Fusion
arXiv:2608.07642v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified re…
arxiv.org1 month agoView details
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
arXiv:2608.07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational…
arxiv.org1 month agoView details
REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs
arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a budget of at most 32B parameters and no model fine-tuning. Our system combines structured chain-of-thought reasoning, rela…
arxiv.org1 month agoView details
arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: e…
arxiv.org1 month agoView details
Protecting patient privacy in clinical foundation models: Technical and legal perspectives
arXiv:2608.07705v1 Announce Type: new Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As deployment expands, privacy risk increasingly arises from model-mediated leakage, yet its prevalence and severity remain poorly qua…
arxiv.org1 month agoView details
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for…
arxiv.org1 month agoView details
Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus
arXiv:2608.10688v1 Announce Type: new Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking h…
arxiv.org1 month agoView details
Data Attribution of Emergent Misalignment with Persona Features
arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligne…
arxiv.org1 month agoView details
arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally…
arxiv.org1 month agoView details
Mitigating Context Interference for Reliable and Efficient Search Agents
arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts of multi-turn search agents are lengthy and complex. For example, the retrieved set of doc…
arxiv.org1 month agoView details
arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma…
arxiv.org1 month agoView details
MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection
arXiv:2608.10459v1 Announce Type: new Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models. Input-only encoder detectors are suitable…
arxiv.org1 month agoView details
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean a…
arxiv.org1 month agoView details
Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a member of the pre-training corpus of a…
arxiv.org1 month agoView details
EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection
arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper prese…
arxiv.org1 month agoView details
arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the…
arxiv.org1 month agoView details
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers: responses may be partly correct, incomplete, or…
arxiv.org1 month agoView details
SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution
arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may expose private user information and incu…
arxiv.org1 month agoView details
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documentation inspection,…
arxiv.org1 month agoView details
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a stable physical object but a latent and temporally evolving attention state. In this work, we…
arxiv.org1 month agoView details
The Illusion of Cross-Lingual Safety in Low-Resource Languages
arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross…
arxiv.org1 month agoView details
JustLLMGRPO: Radiographic Control for Chest X-Ray Generation
arXiv:2608.08046v1 Announce Type: new Abstract: Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain adaptation. We sho…
arxiv.org1 month agoView details
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
arXiv:2608.11171v1 Announce Type: new Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mecha…
arxiv.org1 month agoView details
arXiv:2608.11200v1 Announce Type: new Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while b…
arxiv.org1 month agoView details
arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four…
arxiv.org1 month agoView details
arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. However, existing benchmarks provide limited assessment of whether LLMs can faithfully perform multi-hop reasoning chains a…
arxiv.org1 month agoView details
Back to the Future: A workbook time machine for spread sheet creation benchmarks
arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public workbook corpora,…
arxiv.org1 month agoView details
Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning
arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generation. Tool augmentat…
arxiv.org1 month agoView details
arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent sys…
arxiv.org1 month agoView details
arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse r…
arxiv.org1 month agoView details
Thought-Level Beam Search for Reasoning
arXiv:2608.08020v2 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time…
arxiv.org1 month agoView details
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
arXiv:2608.08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nev…
arxiv.org1 month agoView details
LLM Agents Factory: Retrieval of Domain-Specific LLM Agents
arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user re…
arxiv.org1 month agoView details
PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively under…
arxiv.org1 month agoView details
The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding
arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid masking distorts the model's underlying probability distribution, often biasing generation toward vali…
arxiv.org1 month agoView details
Multimodal Item Parameter Estimation using Simulated Response Probabilitie
arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a l…
arxiv.org1 month agoView details
Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies; in contrast, we estimate corpus ratios…
arxiv.org1 month agoView details
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capa…
arxiv.org1 month agoView details
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we in…
arxiv.org1 month agoView details
arXiv:2608.11008v1 Announce Type: new Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbaggin…
arxiv.org1 month agoView details
arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by nat…
arxiv.org1 month agoView details
arXiv:2608.10627v1 Announce Type: new Abstract: Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims before checking each one. Decomposition itself is treated as a neutral preprocessing step. We show it is not: a decompos…
arxiv.org1 month agoView details
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how th…
arxiv.org1 month agoView details
Attention-Path Fragility as an Uncertainty Signal in Large Language Models
arXiv:2608.11138v1 Announce Type: new Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwor…
arxiv.org1 month agoView details
Divergent Response Modes in Frontier Language Models Under Steering Pressure
arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six…
arxiv.org1 month agoView details
TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents
arXiv:2608.07917v2 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accur…
arxiv.org1 month agoView details
Multiclass Sentiment Analysis for Identifying Political Viewpoints
arXiv:2608.11049v1 Announce Type: new Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) t…
arxiv.org1 month agoView details
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement with CausalNav, a controller built around a signed, action-conditioned transition graph over i…
arxiv.org1 month agoView details
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structu…
arxiv.org1 month agoView details
The Authority Expectancy Effect in Multi-User Conflict
arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemini, GPT, Grok) and t…
arxiv.org1 month agoView details
MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level task co…
arxiv.org1 month agoView details
FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation
arXiv:2608.10916v1 Announce Type: new Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive human-annotated ground truth, or rely on LLM j…
arxiv.org1 month agoView details
The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs
arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bi…
arxiv.org1 month agoView details
CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
arXiv:2608.07965v1 Announce Type: new Abstract: Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecurity education. Generative AI agents offer a path to delivering such experiences adaptively at scale, but introduce well-documented ri…
arxiv.org1 month agoView details
Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final e…
arxiv.org1 month agoView details
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce TelemetrySuffBench, a controlled benchmark that separates failure detection, fault-origin localiza…
arxiv.org1 month agoView details
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
arXiv:2608.07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical r…
arxiv.org1 month agoView details
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions…
arxiv.org1 month agoView details
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the offi…
arxiv.org1 month agoView details
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state p…
arxiv.org1 month agoView details
arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work,…
arxiv.org1 month agoView details
NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
arXiv:2608.07530v1 Announce Type: new Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. Howe…
arxiv.org1 month agoView details
KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs
arXiv:2608.07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different questions. In this work, we…
arxiv.org1 month agoView details
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private…
arxiv.org1 month agoView details
TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair
arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc ed…
arxiv.org1 month agoView details
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. Understanding these relationships is important for model provenance…
arxiv.org1 month agoView details
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
arXiv:2608.07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an…
arxiv.org1 month agoView details
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free qua…
arxiv.org1 month agoView details
Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains
arXiv:2608.07474v1 Announce Type: new Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load.…
arxiv.org1 month agoView details
Assessing Reliability of BERT-Based Models on Question Answering Tasks
arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language processing (NLP), particularly those based on…
arxiv.org1 month agoView details
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps propagate to an in…
arxiv.org1 month agoView details
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
arXiv:2608.07566v1 Announce Type: new Abstract: We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework encodes a scene into coefficients of a fixed symmetric matrix basis, assembles them into a Lie algebra element, and exponentiates the result to a Riemannian or Lor…
arxiv.org1 month agoView details
Controlled Memory Interference in Continual LLM Agents
arXiv:2608.07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise, or interfere with…
arxiv.org1 month agoView details
arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reaso…
arxiv.org1 month agoView details
ReLTEx: Reliable LLM-based Taxonomy Expansion
arXiv:2608.10970v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansions often leads to noi…
arxiv.org1 month agoView details
arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language…
arxiv.org1 month agoView details
arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically lar…
arxiv.org1 month agoView details
Models
View all models →image-text-to-video · minimax-h3 · diffusers · safetensors
huggingface.co1 month ago5400 ptsView details
image-to-video · diffusion-single-file · image-to-video · text-to-video
huggingface.co1 month ago4153 ptsView details
image-text-to-text · transformers · safetensors · muse_glimmer
huggingface.co1 month ago1791 ptsView details
image-to-video · diffusers · t2v · i2v
huggingface.co1 month ago891 ptsView details
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
safetensors · en · arxiv:2410.17196
huggingface.co1 month ago392 ptsView details
image-text-to-video · text-to-video · image-text-to-video · image-to-video
huggingface.co1 month ago345 ptsView details
SexGod1979/PinkCherry_MiniMax-H3
text-to-video · transformers · minimax-h3 · text-to-video
huggingface.co1 month ago337 ptsView details
text-generation · safetensors · bailing_hybrid · text-generation
huggingface.co1 month ago346 ptsView details
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
text-generation · transformers · safetensors · nemotron_h
huggingface.co1 month ago331 ptsView details
drbaph/MiniMax-H3-Turbo-Lora-ComfyUI
text-to-video · minimax-h3 · lora · adapter
huggingface.co1 month ago327 ptsView details
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
text-generation · transformers · safetensors · nemotron_h
huggingface.co1 month ago172 ptsView details
Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot
image-text-to-video · diffusers · text-to-video · image-to-video
huggingface.co1 month ago176 ptsView details
text-generation · cactus-needle · needle · tool-calling
huggingface.co1 month ago161 ptsView details
text-generation · transformers · safetensors · Motif
huggingface.co1 month ago136 ptsView details
SyzygyResearch/Mach-1-Additive-35B
qwen3_5_moe · qwen · mach-1
huggingface.co1 month ago123 ptsView details
peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF
image-text-to-text · gguf · llama.cpp · qwen3_6
huggingface.co1 month ago64 ptsView details
Sachin21112004/distilbart-news-summarizer
summarization · transformers · pytorch · jax
huggingface.co1 month ago21 ptsView details
jaykwok/Qwen3-ASR-1.7B-JA-Anime-Galgame-hf
automatic-speech-recognition · transformers · safetensors · qwen3_asr
huggingface.co1 month ago2 ptsView details
text-generation · transformers · safetensors · llama
huggingface.co1 month agoView details
mobilint/whisper-large-v3-turbo
automatic-speech-recognition · transformers · safetensors · mobilint-whisper
huggingface.co1 month agoView details
naver-ellm/HyperCLOVAX-SEED-Text-Instruct-0.5B-GGUF
gguf · base_model:naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-0.5B · base_model:quantized:naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-0.5B
huggingface.co1 month agoView details
automatic-speech-recognition · transformers · safetensors · mobilint-whisper
huggingface.co1 month agoView details
automatic-speech-recognition · transformers · safetensors · mobilint-whisper
huggingface.co1 month agoView details
Open source
View all open source →<details open> mtmd: support pocket-tts (#26871) * adapt the api * text model ok * working impl, need verify and clean up * mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample was built as one…
github.com1 month agoView details
## [3.0.0](https://github.com/openai/openai-python/compare/v2.54.0...v3.0.0) (2026-08-12) ### ⚠ BREAKING CHANGES * **api:** HTTPX2 is now the default HTTP client, and `httpx` is no longer installed automatically. Applications using custom HTTPX clients, transports, or configurat…
github.com1 month agoView details
<details open> tests : disable backend sampler hip multi output (#26878) * test-backend-sampler: skip multi_output_sampling_chain on HIP The new multi_output_sampling_chain test uses top_k, whose backend probs path needs CUB (unavailable on HIP), so sampled_probs is null and the…
github.com1 month agoView details
langchain-ai/langchain langchain-anthropic==1.5.5
Changes since langchain-anthropic==1.5.4 release(anthropic): 1.5.5 (#39597) fix(anthropic): report reasoning tokens in usage metadata (#39590) fix(anthropic): fix `KeyError` on rename in Claude file-tool middleware (#39293) chore(model-profiles): refresh model profile data (#392…
github.com1 month agoView details
langchain-ai/langchain langchain==1.3.15
Changes since langchain==1.3.14 release(langchain): 1.3.15 (#39595) feat(langchain): expose `trace_policy` on `AgentMiddleware` (#38910) chore(langchain): fix type errors in tests (#39589) chore: bump h2 from 4.3.0 to 4.4.1 in /libs/langchain_v1 (#39324) fix(langchain): preserve…
github.com1 month agoView details
## [2.54.0](https://github.com/openai/openai-python/compare/v2.53.0...v2.54.0) (2026-08-11) ### Features * **api:** Add new Responses model identifiers ([#3595](https://github.com/openai/openai-python/issues/3595)) ([0652787](https://github.com/openai/openai-python/commit/065278…
github.com1 month agoView details
langchain-ai/langchain langchain-core==1.5.4
Changes since langchain-core==1.5.3 release(core): 1.5.4 (#39592) fix(core): compat with pydantic 2.14 (#39328) fix(core): stop StructuredPrompt from mutating caller kwargs (#39174) fix(core): preserve flat tool args schema for `RootModel` runnables (#39307) fix(core): close int…
github.com1 month agoView details
<details open> model : fix SWA not being enabled for EXAONE 4.5 (#26848) * model : fix SWA not being enabled for EXAONE 4.5 load_arch_hparams tests `hparams.n_layer() == 64` before LLM_KV_NEXTN_PREDICT_LAYERS has been read. n_layer() returns n_layer_all - n_layer_nextn and n_lay…
github.com1 month agoView details
This is a patch release on top of v0.27.0. - Support quantized DSpark Markov heads (#50424)
github.com1 month agoView details