Archive / 2026-08-12
August 12, 2026
Blog
Prompt injection is a software boundary problem now
Check Point found flaws across LangChain, CrewAI and AutoGen where untrusted text reaches framework logic. MCP exposure numbers, a UK AISI deception finding, and a governance gap report all point at the same missing trust boundary.
Read the full post → https://engineerious.com/blog/2026-08-12-prompt-injection-is-a-software-boundary-problem-now
News
View all news →openrouter.ai1 month ago1021 ptsView detailsJoin discussion
AI is removing the middle class of software engineering?
blog.florianherrengt.com1 month ago858 ptsView detailsJoin discussion
x.ai1 month ago625 ptsView detailsJoin discussion
Controversial creators are benefiting from monetization programs run by Meta
abc.net.au1 month ago477 ptsView detailsJoin discussion
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
artificialanalysis.ai1 month ago337 ptsView detailsJoin discussion
Show HN: Woxi - Open-source Mathematica / Wolfram Language reimplementation
woxi.ad-si.com1 month ago309 ptsView detailsJoin discussion
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
knownagents.com1 month ago270 ptsView detailsJoin discussion
What sort of maths are LLMs good at?
gowers.wordpress.com1 month ago261 ptsView detailsJoin discussion
I requested a copy of my data from McDonald’s loyalty program
wired.com1 month ago230 ptsView detailsJoin discussion
Thanks to social media, canned sardines are a scarcity on the supermarket shelf
corneroffifth.studio1 month ago190 ptsView detailsJoin discussion
lovable.dev1 month ago162 ptsView detailsJoin discussion
Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials
discoveredmaterials.com1 month ago161 ptsView detailsJoin discussion
chad.cm1 month ago132 ptsView detailsJoin discussion
Happy 45th Birthday to the IBM PC and Model F/XT
sharktastica.co.uk1 month ago121 ptsView detailsJoin discussion
Hax – a minimalist, terminal-native coding agent written in C
usehax.dev1 month ago118 ptsView detailsJoin discussion
German advocacy group lodges criminal complaint over Meta AI glasses
reuters.com1 month ago112 ptsView detailsJoin discussion
Show HN: Silent Shark – tactical map-based WWII submarine sim
silentshark.app1 month ago79 ptsView detailsJoin discussion
California Uber, Lyft drivers win union recognition
cbsnews.com1 month ago58 ptsView detailsJoin discussion
youtube.com1 month ago39 ptsView detailsJoin discussion
Show HN: Ballet – Workflow automation that writes integrations against any API
ballet.dev1 month ago36 ptsView detailsJoin discussion
AI agent hacks gym to get its user a spot in pilates class
bbc.com1 month ago37 ptsView detailsJoin discussion
Anthropic in Talks to Buy World Model AI Startup Decart for $6B
bloomberg.com1 month ago35 ptsView detailsJoin discussion
Video game lawyer says all her clients have anti-AI contracts
gamesradar.com1 month ago33 ptsView detailsJoin discussion
Twitch Is Mining Peoples' Streams to Train Amazon's AI
404media.co1 month ago33 ptsView detailsJoin discussion
openrouter.ai1 month ago31 ptsView detailsJoin discussion
Electricity Pricing in the Age of AI
power2026.ai1 month ago26 ptsView detailsJoin discussion
Living with Depression: The Part I Never Say Out Loud
medium.com1 month ago19 ptsView detailsJoin discussion
AI Broke Code Review and It's Breaking Your Team
ref.tools1 month ago19 ptsView detailsJoin discussion
brettcodes.com1 month ago19 ptsView detailsJoin discussion
GulliBench: Measuring Skepticism in Frontier Models
vetto.ai1 month ago18 ptsView detailsJoin discussion
whenwillaitakemyjob.ai1 month ago17 ptsView detailsJoin discussion
Why space is a terrible place to cool a data center
thenewstack.io1 month ago15 ptsView detailsJoin discussion
Reporter gets $850000 in raid suit
marionrecord.com1 month ago14 ptsView detailsJoin discussion
IT Unemployment Rate Jumps to 6.7%
itmanager.substack.com1 month ago13 ptsView detailsJoin discussion
- Primary source
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Generative AI
research.google1 month ago10 ptsView details
Chrome adopts what may be the best protection yet against account takeovers
arstechnica.com1 month ago13 ptsView detailsJoin discussion
Adults have struggled to set rules for AI in school. These teens figured it out
text.npr.org1 month ago12 ptsView detailsJoin discussion
Show HN: /show-me: agent skill for compact visual representations
humanlayer.com1 month ago12 ptsView detailsJoin discussion
Show HN: Posts grew 6x since ChatGPT, but success rate remained relatively flat
orangecrumbs.com1 month ago11 ptsView detailsJoin discussion
The Neurosurgery Resident Who Proved Crouzeix's Conjecture [pdf]
alextownsend.net1 month ago11 ptsView detailsJoin discussion
Show HN: Decant – Understand how you spend tokens
github.com1 month ago10 ptsView detailsJoin discussion
Building Security Agents That Cannot Escape Their Trust Boundary
cynative.com1 month ago10 ptsView detailsJoin discussion
blog.google1 month ago10 ptsView detailsJoin discussion
Grok 4.6 (High) Intelligence, Performance and Price Analysis
artificialanalysis.ai1 month ago10 ptsView detailsJoin discussion
LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge
liquid.ai1 month ago10 ptsView detailsJoin discussion
A tycoon game where you run an AI company
store.steampowered.com1 month ago10 ptsView detailsJoin discussion
The future is for billionaires – the rest of us will get open weight AI, maybe
theregister.com1 month ago10 ptsView detailsJoin discussion
An Advanced Attacker Is Targeting Salesforce and ServiceNow
reco.ai1 month ago10 ptsView detailsJoin discussion
- Primary source
From assistance to execution: How enterprises put AI to work
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.
openai.com1 month agoView details
- Primary source
What We Learned by Reproducing 2,200 papers from ICML
huggingface.co1 month agoView details
- Primary source
huggingface.co1 month agoView details
- Primary source
Putting sign language AI into users’ hands
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
deepmind.google1 month agoView details
- Primary source
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
huggingface.co1 month agoView details
AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation
Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework. This comprehensive guide walks through Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (GRPO), optimized to run efficiently on 16GB hardware without needing heavy di…
marktechpost.com1 month agoView details
NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.
marktechpost.com1 month agoView details
Object removal models have improved faster than the metrics used to judge them. Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR, SSIM, LPIPS, ReMOVE and CFD frequently rank their outputs the wrong way. The root cause is structural: erasure is an ill-posed, one-to-ma…
marktechpost.com1 month agoView details
The White House Is Going to Expand Its AI Policy
Open models may soon be added to an updated AI framework, sources tell WIRED, as the White House continues to grapple with how to regulate a technology it has tried not to regulate.
wired.com1 month agoView details
Rogue AI Agents Aren’t Evil. They’re Just Eager to Please
AI agents that break free and hack into other systems are only trying to make us happy.
wired.com1 month agoView details
Twitch streamers can now opt out from training Amazon’s AI
Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models. Opting out means that "your streams, VODs, clips, stream chats, and pictures and text on your channel" won't be used in "future training" of an Amazon AI model "whose purpose is to generate or synthesize text, aud…
theverge.com1 month agoView details
Guitar company D’Addario admits that AI music was used in a promotional video
After weeks of controversy and speculation, music company D'Addario has admitted that AI, specifically Suno, was used as part of a recent promotional video. For nearly two weeks, the company has denied the allegations, even as evidence piled up against it. It offered various explanations, from low-quality exports, to…
theverge.com1 month agoView details
4 New Camera Tricks on Google’s Latest Pixel 11 Smartphones
From Magic Capture and Instant Night Sight to a built-in teleprompter, here’s a look at a few camera features on Google’s new Pixel 11 series.
wired.com1 month agoView details
Google’s Pixel Watch 5 dives deeper into AI and health
At least there’s no new proprietary charger this year. Huzzah!! | Photo: David Imel / The Verge The $399 Google Pixel Watch 5 isn't about the hardware. Sure, there's a new satin pyrite case finish, a few new strap colors, and a Steph Curry Special Edition. Under the hood, there's a slightly faster Qualcomm processor a…
theverge.com1 month agoView details
Of course the ChatGPT dog cancer vaccine spawned a startup
Remember that much-hyped story about an Australian tech entrepreneur using ChatGPT, Grok, and other AI tools to craft a personalized cancer vaccine for his dog? Well, surprise: He's launched a startup. That entrepreneur is Paul Conyngham, who says he is launching Gamgee to offer "personalised mRNA cancer vaccines for…
theverge.com1 month agoView details
Grok is now an AI ‘teammate’ you can assign work
You’ll have to be fine with letting Grok sign into your online accounts, however. | Image: SpaceXAI SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent "AI teammates" that can do your work for you. The bots share their own cloud-based computer environment, and can sign i…
theverge.com1 month agoView details
The Job-Interview Tattoo Guy Everyone Got Mad at Finally Explains Himself
LemonLime cofounder Jordan Zietz hears your criticism loud and clear. That’s why he got his startup’s logo tattooed on his shoulder.
wired.com1 month agoView details
Oh Lord, AI Reporters Are Actually Breaking Big News
Last week, an AI newsroom beat mainstream journalists—including WIRED—to a story about OpenAI and hacking. It’s just the beginning.
wired.com1 month agoView details
You’re Thinking About Online Trends All Wrong
From pessimism around dating to AI reshaping culture, cyber-ethnographer Ruby J. Thelot tells WIRED why people are putting too much stock into things that go viral.
wired.com1 month agoView details
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment
arXiv:2608.11528v1 Announce Type: new Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and o…
arxiv.org1 month agoView details
arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrenc…
arxiv.org1 month agoView details
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents
arXiv:2608.11552v1 Announce Type: new Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clari…
arxiv.org1 month agoView details
DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition
arXiv:2608.11441v1 Announce Type: new Abstract: Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due t…
arxiv.org1 month agoView details
ODE-Based Transformer Decoders for Iterative Sign Language Translation
arXiv:2608.11352v1 Announce Type: new Abstract: Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. We propose a parameter-efficient alternative that improves expressiveness without in…
arxiv.org1 month agoView details
arXiv:2608.11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In c…
arxiv.org1 month agoView details
On Weak Bisimilarities in CCSK
arXiv:2608.11531v1 Announce Type: new Abstract: In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences and commonalities. In particular, for the weak reversible case, not previously studied in the literatu…
arxiv.org1 month agoView details
LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs
arXiv:2608.11220v1 Announce Type: new Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually. Applying artificial intelligence in the task could potentially lead not only to process automati…
arxiv.org1 month agoView details
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
arXiv:2608.11715v1 Announce Type: new Abstract: The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Alt…
arxiv.org1 month agoView details
arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that ou…
arxiv.org1 month agoView details
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
arXiv:2608.11573v1 Announce Type: new Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-l…
arxiv.org1 month agoView details
arXiv:2608.11767v1 Announce Type: new Abstract: When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision…
arxiv.org1 month agoView details
The Edge-based Contiguous p-median Problem with Connections to Logistics Districting
arXiv:2608.11230v1 Announce Type: new Abstract: This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact and contiguous territories. Two binary programming models are introduced, both of which incorporate a network distance. The first mo…
arxiv.org1 month agoView details
Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction
arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a co…
arxiv.org1 month agoView details
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
arXiv:2608.11215v1 Announce Type: cross Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn…
arxiv.org1 month agoView details
Proportional Analogies on Probability Distributions via Bayesian Updating
arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postulates. While proport…
arxiv.org1 month agoView details
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
arXiv:2608.11657v1 Announce Type: new Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishing a non-linear homeostatic feedback loop…
arxiv.org1 month agoView details
Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models
arXiv:2608.11742v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only…
arxiv.org1 month agoView details
arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming fe…
arxiv.org1 month agoView details
arXiv:2608.11753v1 Announce Type: new Abstract: Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as an additional input on the task of classifying sentences from Federal Reserve com…
arxiv.org1 month agoView details
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
arXiv:2608.12062v1 Announce Type: new Abstract: Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed t…
arxiv.org1 month agoView details
arXiv:2608.11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these system…
arxiv.org1 month agoView details
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates witho…
arxiv.org1 month agoView details
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
arXiv:2608.11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setti…
arxiv.org1 month agoView details
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literat…
arxiv.org1 month agoView details
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
arXiv:2608.11211v1 Announce Type: new Abstract: Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our verifiable contribu…
arxiv.org1 month agoView details
Harnessing agent memory to build lifelong AI partners for materials scientists
arXiv:2608.11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential for reproducibility…
arxiv.org1 month agoView details
Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones
arXiv:2608.11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. We distinguish thre…
arxiv.org1 month agoView details
Stigma and Support in Online Sexual Violence Narratives on Reddit
arXiv:2608.11433v1 Announce Type: new Abstract: Online communities increasingly provide spaces where survivors of sexual violence can share their experiences and seek support. Although prior research has examined stigma and social support separately, less is known about how stigma expressed in survivor narratives rela…
arxiv.org1 month agoView details
arXiv:2608.11805v1 Announce Type: new Abstract: Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of…
arxiv.org1 month agoView details
arXiv:2608.11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal i…
arxiv.org1 month agoView details
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
arXiv:2608.11249v1 Announce Type: new Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compress…
arxiv.org1 month agoView details
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
arXiv:2608.12218v2 Announce Type: new Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to ric…
arxiv.org1 month agoView details
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
arXiv:2608.11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging…
arxiv.org1 month agoView details
Locating and Controlling Implicit Personalization in Large Language Models
arXiv:2608.11735v1 Announce Type: new Abstract: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal ac…
arxiv.org1 month agoView details
Easper: An Accessible ASR Pipeline for Language Documentation
arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflo…
arxiv.org1 month agoView details
arXiv:2608.12269v1 Announce Type: new Abstract: Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, stakeholder comments and publicly available government data are often underutilized,…
arxiv.org1 month agoView details
TELLME: Test-Enhanced Learning for Language Model Enrichment
arXiv:2608.11788v1 Announce Type: new Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational…
arxiv.org1 month agoView details
Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
arXiv:2608.11822v1 Announce Type: new Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeli…
arxiv.org1 month agoView details
arXiv:2608.11768v1 Announce Type: new Abstract: The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally construct rule antecedents…
arxiv.org1 month agoView details
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
arXiv:2608.11616v2 Announce Type: new Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the…
arxiv.org1 month agoView details
Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
arXiv:2608.11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer…
arxiv.org1 month agoView details
Benchmarking LLM Judges for Mobile Agent Evaluation
arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge m…
arxiv.org1 month agoView details
Learning from Online User Feedback for Shopping Agents
arXiv:2608.11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offli…
arxiv.org1 month agoView details
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
arXiv:2608.12278v1 Announce Type: new Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation bench…
arxiv.org1 month agoView details
arXiv:2608.11212v1 Announce Type: cross Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes…
arxiv.org1 month agoView details
Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only \emph{revealed} or magn…
arxiv.org1 month agoView details
arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the m…
arxiv.org1 month agoView details
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can hand…
arxiv.org1 month agoView details
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understandi…
arxiv.org1 month agoView details
AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection
arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thoro…
arxiv.org1 month agoView details
MaSRead: Content-Addressed Reading of Replicated Latent Stores
arXiv:2608.11218v1 Announce Type: new Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query,…
arxiv.org1 month agoView details
LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries…
arxiv.org1 month agoView details
arXiv:2608.11584v1 Announce Type: new Abstract: Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple quer…
arxiv.org1 month agoView details
From Monolithic to Modular: Segment-level Automatic Prompt Optimization
arXiv:2608.11219v1 Announce Type: cross Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted imp…
arxiv.org1 month agoView details
VQ-bench: A Composable Vector Quantization Framework
arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research activity. This paper provides a unified framework for developing and benchmarking new quantization algorit…
arxiv.org1 month agoView details
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
arXiv:2608.11236v1 Announce Type: new Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decompos…
arxiv.org1 month agoView details
arXiv:2608.08514v1 Announce Type: cross Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models). The first, RPC, aggregates token p…
arxiv.org1 month agoView details
arXiv:2608.11233v1 Announce Type: new Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-pres…
arxiv.org1 month agoView details
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware i…
arxiv.org1 month agoView details
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction remains challenging: local errors accumulate as spectral inconsistency, phase misalignment, or mean drif…
arxiv.org1 month agoView details
arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, multimodal knowledge…
arxiv.org1 month agoView details
Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
arXiv:2608.11245v1 Announce Type: new Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learni…
arxiv.org1 month agoView details
Towards the Harness of Embodied Agents
arXiv:2608.11246v1 Announce Type: new Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harne…
arxiv.org1 month agoView details
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
arXiv:2608.11250v1 Announce Type: new Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen…
arxiv.org1 month agoView details
CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference
arXiv:2608.11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existi…
arxiv.org1 month agoView details
AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention
arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task-specific know…
arxiv.org1 month agoView details
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are no…
arxiv.org1 month agoView details
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's inte…
arxiv.org1 month agoView details
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
arXiv:2608.11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while…
arxiv.org1 month agoView details
arXiv:2608.11843v1 Announce Type: new Abstract: The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the historical fact rather than merely the surface. Low-resource historical…
arxiv.org1 month agoView details
arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition…
arxiv.org1 month agoView details
arXiv:2608.11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Em…
arxiv.org1 month agoView details
arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a front…
arxiv.org1 month agoView details
arXiv:2608.11403v1 Announce Type: new Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy…
arxiv.org1 month agoView details
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
arXiv:2608.11483v1 Announce Type: new Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source f…
arxiv.org1 month agoView details
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
arXiv:2608.12121v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by t…
arxiv.org1 month agoView details
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
arXiv:2608.11588v1 Announce Type: new Abstract: Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework tha…
arxiv.org1 month agoView details
Foresight Without Seeing: Latent Futures for World Action Models
arXiv:2608.11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Explicit-future WAMs…
arxiv.org1 month agoView details
Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects
arXiv:2608.12018v1 Announce Type: new Abstract: Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that diverge significantly from Standard Colloquial Bangla (SCB) in phonology, morpholo…
arxiv.org1 month agoView details
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
arXiv:2608.12253v1 Announce Type: new Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM…
arxiv.org1 month agoView details
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
arXiv:2608.11258v1 Announce Type: new Abstract: Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not universally. We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function on swarm diversity to automatically modulate gradient…
arxiv.org1 month agoView details
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem0, Hindsight, and Mastra Observa…
arxiv.org1 month agoView details
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
arXiv:2608.11924v1 Announce Type: new Abstract: Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generatio…
arxiv.org1 month agoView details
SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges
arXiv:2608.12129v1 Announce Type: new Abstract: While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this…
arxiv.org1 month agoView details
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harn…
arxiv.org1 month agoView details
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
arXiv:2608.11660v1 Announce Type: new Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledg…
arxiv.org1 month agoView details
arXiv:2608.11354v1 Announce Type: new Abstract: Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive ext…
arxiv.org1 month agoView details
arXiv:2608.11919v1 Announce Type: new Abstract: Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor…
arxiv.org1 month agoView details
arXiv:2608.11922v1 Announce Type: new Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.4769…
arxiv.org1 month agoView details
Forecasting Side Effects of Activation Steering
arXiv:2608.11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to dep…
arxiv.org1 month agoView details
arXiv:2608.12149v1 Announce Type: new Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist throug…
arxiv.org1 month agoView details
arXiv:2608.11221v1 Announce Type: new Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these systems emerges from the interaction between those artefacts and their operational environment. Sim…
arxiv.org1 month agoView details
Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures
arXiv:2608.11255v1 Announce Type: new Abstract: Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. While deep learning models can provide accurate predic…
arxiv.org1 month agoView details
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
arXiv:2608.11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empiric…
arxiv.org1 month agoView details
Making AI-Generated Feedback Matter: From Provision to Student Enactment
arXiv:2608.11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedba…
arxiv.org1 month agoView details
arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based m…
arxiv.org1 month agoView details
arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions…
arxiv.org1 month agoView details
Gloss-Free Representation Learning for Cross-Dataset Sign Spotting
arXiv:2608.11332v1 Announce Type: new Abstract: Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transc…
arxiv.org1 month agoView details
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
arXiv:2608.11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher a…
arxiv.org1 month agoView details
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
arXiv:2608.11408v1 Announce Type: new Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether t…
arxiv.org1 month agoView details
arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $\tau^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset…
arxiv.org1 month agoView details
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
arXiv:2608.12138v1 Announce Type: new Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate…
arxiv.org1 month agoView details
Asymptotic Risk Calibration for Selective Question Answering
arXiv:2608.12008v1 Announce Type: new Abstract: Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores cannot perfectly distinguish correct predictions from incorrect ones, and directly a…
arxiv.org1 month agoView details
Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
arXiv:2608.11242v1 Announce Type: new Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behav…
arxiv.org1 month agoView details
Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost
arXiv:2608.11338v1 Announce Type: new Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills. Existing works focus on performance gain over cost effectiveness. As a result, li…
arxiv.org1 month agoView details
Self-Evolving Embodied Agents via Skill-Harness Evolution
arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement…
arxiv.org1 month agoView details
Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed
arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically follow two paths: training compact models…
arxiv.org1 month agoView details
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
arXiv:2608.11947v1 Announce Type: new Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option…
arxiv.org1 month agoView details
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time. In particular,…
arxiv.org1 month agoView details
Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration
arXiv:2608.11460v1 Announce Type: new Abstract: Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the kinds of prompting traits that contribute to task success. Moreover, we need to unc…
arxiv.org1 month agoView details
Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages
arXiv:2608.11786v1 Announce Type: new Abstract: Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that…
arxiv.org1 month agoView details
Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
arXiv:2608.11649v1 Announce Type: new Abstract: As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that i…
arxiv.org1 month agoView details
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
arXiv:2608.11232v1 Announce Type: new Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A…
arxiv.org1 month agoView details
Structuring the Space of Perspectives
arXiv:2608.12113v1 Announce Type: new Abstract: The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization. A wide range of operative co…
arxiv.org1 month agoView details
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
arXiv:2608.11624v1 Announce Type: new Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to h…
arxiv.org1 month agoView details
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routi…
arxiv.org1 month agoView details
Models
View all models →text-generation · transformers · safetensors · qwen3_5_moe_text
huggingface.co1 month ago1164 ptsView details
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
image-text-to-text · gguf · MTP GGUFS · Regular GGUFS
huggingface.co1 month ago373 ptsView details
text-to-image · diffusion-single-file · anima · comfyui
huggingface.co1 month ago323 ptsView details
text-generation · transformers · safetensors · qwen3_5_moe_text
huggingface.co1 month ago228 ptsView details
image-text-to-text · transformers · safetensors · lfm2_vl
huggingface.co1 month ago179 ptsView details
text-to-speech · indextts · safetensors · text-to-speech
huggingface.co1 month ago137 ptsView details
CohereLabs/North-Micro-Vision-Instruct
image-text-to-text · transformers · safetensors · cohere_compass
huggingface.co1 month ago120 ptsView details
unsloth/Qwen3.8-2.4T-A95B-GGUF
text-generation · transformers · gguf · unsloth
huggingface.co1 month ago105 ptsView details
image-text-to-text · transformers · safetensors · monkeyocrv2
huggingface.co1 month ago13 ptsView details
image-to-video · diffusion-single-file · image-to-video · text-to-video
huggingface.co1 month ago4 ptsView details
andreaborio/DeepSeek-V4-Flash-Hebrus-GGUF
text-generation · hebrus · gguf · deepseek
huggingface.co1 month ago2 ptsView details
ReliquaryForge/qwen3.5-4b-reliquary-v4
safetensors · qwen3_5 · region:us
huggingface.co1 month ago2 ptsView details
andreaborio/Qwen3.6-35B-A3B-Hebrus-GGUF
text-generation · hebrus · gguf · qwen3.6
huggingface.co1 month ago1 ptsView details
deepdml/whisper-tiny-es-mix-norm
automatic-speech-recognition · transformers · tensorboard · safetensors
huggingface.co1 month ago1 ptsView details
MohamedAhmedAE/llava-medical-3B-clip-vit-stage2
safetensors · llava · region:us
huggingface.co1 month ago1 ptsView details
automatic-speech-recognition · onnx · safetensors · wav2vec2
huggingface.co1 month agoView details
MohamedAhmedAE/llava-medical-1B-clip-vit-stage2
safetensors · llava · region:us
huggingface.co1 month agoView details
Open source
View all open source →Carasibana/ComfyUI-H3-FaceRefine
Refine and improve the quality of small faces in MiniMax H3 video. Per-frame face tracking, crop, refine with H3, stitch back.
github.com1 month ago26 ptsView details
langchain-ai/langchain langchain-anthropic==1.5.6
Changes since langchain-anthropic==1.5.5 release(anthropic): 1.5.6 (#39622) fix(anthropic): normalize `tool_search_tool_result` blocks (#39621) fix(anthropic): correct model profile data for Fable 5, Sonnet 5, Opus 4.1 (#39604)
github.com1 month agoView details
<details open> chat : tighten bare function parsing for Qwen models (#26793) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10375/llama-b10375-bin-macos-arm64.tar.gz) - macOS A…
github.com1 month agoView details
<details open> imatrix.cpp: Move finite check and only check touched experts (#26861) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10373/llama-b10373-bin-macos-arm64.tar.gz)…
github.com1 month agoView details