Skip to content

Archive / 2026-08-14

August 14, 2026

  1. AI by Hand

    byhand.ai1 month ago357 ptsView detailsJoin discussion

  2. DeepSeek peak/off-peak pricing update

    api-docs.deepseek.com1 month ago238 ptsView detailsJoin discussion

  3. Dear people who work at the airport

    life-after-ssri.bearblog.dev1 month ago208 ptsView detailsJoin discussion

  4. Racket v9.3

    blog.racket-lang.org1 month ago141 ptsView detailsJoin discussion

  5. A simple fix for LLM tail latency

    engineering.myhoai.com1 month ago53 ptsView detailsJoin discussion

  6. Open WireGuard Endpoints

    proxylity.com1 month ago45 ptsView detailsJoin discussion

  7. Baking a Model: A Metaphor for LLM Training

    newsletter.kentbeck.com1 month ago32 ptsView detailsJoin discussion

  8. These ‘Masturbation Consultants’ Were Hired to Pleasure Themselves With AI

    Joi AI hired 10 people to masturbate using AI companions as part of a monthlong “wellness” study. The company claims the practice could help “solve male loneliness.”

    wired.com1 month ago12 ptsView details

  9. Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

    Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training. Terminal-Bench 3.0 moves from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity mov…

    marktechpost.com1 month agoView details

  10. Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

    Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU. The post Meet Needle 2: An Open 45M-Pa…

    marktechpost.com1 month agoView details

  11. The Next Big Influencer Is This 4-Foot-Tall Robot From China

    The Unitree G1 has found online fame as a relatively affordable robot that can charm a crowd. But can it ever hold down a real job?

    wired.com1 month agoView details

  12. Mark Zuckerberg has an Instagzam

    Instagram's wordmark is iconic. Well, was iconic. Apparently Instagram thought it looked old, so the company rolled out a new one this week. It doesn't look like the old Instagram wordmark. It doesn't even look like it spells Instagram anymore. And we cannot figure out why Instagram decided to do this. On this episode…

    theverge.com1 month agoView details

  13. You can now turn off Google Gemini’s visible watermarks

    Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-righ…

    theverge.com1 month agoView details

  14. Tech Visionary Says the Big AI Labs Don’t Get What People Want

    Tim O’Reilly built a publishing empire that AI is helping to destroy. Yet he loves AI—as long as it’s open source.

    wired.com1 month agoView details

  15. People Are ‘Marrying’ Chatbots. These Lawmakers Want to Stop Them

    Human-AI marriages are not currently recognized by US law. Some Republican state policymakers are drafting legislation to keep it that way.

    wired.com1 month agoView details

  16. Apple trained its own AI model for China with help from Alibaba

    Apple has reportedly trained a custom AI model for the China market alongside domestic tech giant Alibaba, a rare cross-border partnership that cuts across growing tensions between Beijing and Washington. The China-focused large language model was developed in partnership with Alibaba and trained with the company's su…

    theverge.com1 month agoView details

  17. EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

    arXiv:2608.13072v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. E…

    arxiv.org1 month agoView details

  18. General Probabilities of Causation with Causal Knowledge

    arXiv:2608.12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the probability of neces…

    arxiv.org1 month agoView details

  19. Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

    arXiv:2608.12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus…

    arxiv.org1 month agoView details

  20. VALG: An Agentic System for ML Theory Research

    arXiv:2608.13060v1 Announce Type: new Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires th…

    arxiv.org1 month agoView details

  21. Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

    arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming.…

    arxiv.org1 month agoView details

  22. From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

    arXiv:2608.13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being…

    arxiv.org1 month agoView details

  23. Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic

    arXiv:2608.13018v1 Announce Type: new Abstract: Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions. In this paper, we…

    arxiv.org1 month agoView details

  24. OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways

    arXiv:2608.12995v1 Announce Type: new Abstract: Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a…

    arxiv.org1 month agoView details

  25. Position: Reasoning is a Learnable Rule-Based Process

    arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and rapid progress, the…

    arxiv.org1 month agoView details

  26. Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

    arXiv:2608.13108v1 Announce Type: new Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate ev…

    arxiv.org1 month agoView details

  27. Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment

    arXiv:2608.13100v1 Announce Type: new Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and au…

    arxiv.org1 month agoView details

  28. @skills: Attention is all you have

    arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This leaves the long tai…

    arxiv.org1 month agoView details

  29. Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

    arXiv:2608.13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidenc…

    arxiv.org1 month agoView details

  30. DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

    arXiv:2608.13048v1 Announce Type: new Abstract: In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patter…

    arxiv.org1 month agoView details

  31. Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

    arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appeal…

    arxiv.org1 month agoView details

  32. DiG-bench: Discovery in Games

    arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation…

    arxiv.org1 month agoView details

  33. Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

    arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifac…

    arxiv.org1 month agoView details

  34. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

    arXiv:2608.13120v1 Announce Type: new Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback…

    arxiv.org1 month agoView details

  35. Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

    arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to…

    arxiv.org1 month agoView details

  36. Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

    arXiv:2608.12373v1 Announce Type: new Abstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stak…

    arxiv.org1 month agoView details

  37. Designing AI Pipelines for Decision-Ready ITSM Intelligence

    arXiv:2608.12670v1 Announce Type: new Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and evaluated following…

    arxiv.org1 month agoView details

  38. Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

    arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushbac…

    arxiv.org1 month agoView details

  39. Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy

    arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across item variants disto…

    arxiv.org1 month agoView details

  40. Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs

    arXiv:2608.12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG has focused on preventing unauthorized users from accessing sensitive data. However, another importa…

    arxiv.org1 month agoView details

  41. The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

    arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature ex…

    arxiv.org1 month agoView details

  42. Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

    arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by incr…

    arxiv.org1 month agoView details

  43. Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

    arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fi…

    arxiv.org1 month agoView details

  44. ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    arXiv:2608.12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Co…

    arxiv.org1 month agoView details

  45. CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

    arXiv:2608.12842v1 Announce Type: new Abstract: Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade merged-model performa…

    arxiv.org1 month agoView details

  46. Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

    arXiv:2608.12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for…

    arxiv.org1 month agoView details

  47. Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

    arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories in…

    arxiv.org1 month agoView details

  48. Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

    arXiv:2608.12895v1 Announce Type: new Abstract: Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0…

    arxiv.org1 month agoView details

  49. Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$

    arXiv:2608.12961v1 Announce Type: new Abstract: The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has not been investigated…

    arxiv.org1 month agoView details

  50. $\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution

    arXiv:2608.12522v1 Announce Type: new Abstract: LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEvo, a framework for…

    arxiv.org1 month agoView details

  51. Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

    arXiv:2608.12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanni…

    arxiv.org1 month agoView details

  52. Uniform Herding: Exemplar Replay with Representation Refresh

    arXiv:2608.13061v1 Announce Type: new Abstract: As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to r…

    arxiv.org1 month agoView details

  53. CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

    arXiv:2608.12555v1 Announce Type: new Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an ident…

    arxiv.org1 month agoView details

  54. Trie Automata for Constrained Decoding over Large Finite Sets

    arXiv:2608.12574v1 Announce Type: new Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar comp…

    arxiv.org1 month agoView details

  55. PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

    arXiv:2608.12762v1 Announce Type: new Abstract: Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manual…

    arxiv.org1 month agoView details

  56. Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

    arXiv:2608.12935v1 Announce Type: new Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppo…

    arxiv.org1 month agoView details

  57. MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

    arXiv:2608.12428v1 Announce Type: new Abstract: Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, o…

    arxiv.org1 month agoView details

  58. Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

    arXiv:2608.13129v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerica…

    arxiv.org1 month agoView details

  59. Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

    arXiv:2608.12476v1 Announce Type: new Abstract: Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitempor…

    arxiv.org1 month agoView details

  60. Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

    arXiv:2608.12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where…

    arxiv.org1 month agoView details

  61. Correct Is Not Governed: Provenance Integrity in Agentic Workflows

    arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define go…

    arxiv.org1 month agoView details

  62. AI and Consumer Rights in India Working Paper

    arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services,…

    arxiv.org1 month agoView details

  63. Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

    arXiv:2608.12346v1 Announce Type: new Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibil…

    arxiv.org1 month agoView details

  64. FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

    arXiv:2608.12932v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Vis…

    arxiv.org1 month agoView details

  65. Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing

    arXiv:2608.12371v1 Announce Type: new Abstract: Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes \emph{MAS-…

    arxiv.org1 month agoView details

  66. Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

    arXiv:2608.12385v1 Announce Type: new Abstract: As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregres…

    arxiv.org1 month agoView details

  67. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

    arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluat…

    arxiv.org1 month agoView details

  68. Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

    arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive…

    arxiv.org1 month agoView details

  69. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

    arXiv:2608.12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or…

    arxiv.org1 month agoView details

  70. Research Assistant: AstraZeneca's Agentic System for R&D

    arXiv:2608.12395v1 Announce Type: new Abstract: We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scient…

    arxiv.org1 month agoView details

  71. On the Expressive Power of Transformers

    arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expressive power of tra…

    arxiv.org1 month agoView details

  72. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

    arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an inciden…

    arxiv.org1 month agoView details

  73. BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs

    arXiv:2608.13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final r…

    arxiv.org1 month agoView details

  74. Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

    arXiv:2608.12892v1 Announce Type: new Abstract: Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path a…

    arxiv.org1 month agoView details

  75. ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

    arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to decompose fact verific…

    arxiv.org1 month agoView details

  76. SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

    arXiv:2608.13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy.…

    arxiv.org1 month agoView details

  1. Qwen/Qwen3.8-27B

    image-text-to-text · transformers · safetensors · qwen3_5

    huggingface.co1 month ago15458 ptsView details

  2. Qwen/Qwen3.8-27B-FP8

    image-text-to-text · transformers · safetensors · qwen3_5

    huggingface.co1 month ago839 ptsView details

  3. FrontiersMind/Lumma-0.6B-Base

    text-generation · transformers · safetensors · Lumma

    huggingface.co1 month ago19 ptsView details

  4. syvai/hviske-v5.3

    automatic-speech-recognition · transformers · safetensors · cohere_asr

    huggingface.co1 month ago10 ptsView details

  5. 0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUF

    image-text-to-text · gguf · llama.cpp · quantization

    huggingface.co1 month ago3 ptsView details

  6. ForeverBlue/Qwen3-VL-2B-GRACE-BF16

    image-text-to-text · transformers · safetensors · qwen3_vl

    huggingface.co1 month ago1 ptsView details

  7. liuff1568/MyAwesomeModel-TestRepo

    feature-extraction · transformers · pytorch · bert

    huggingface.co1 month agoView details

  8. proyeksigila/languagemodel

    gguf · endpoints_compatible · region:us

    huggingface.co1 month agoView details

  9. jcoholich/your_repo_id

    robotics · lerobot · safetensors · robotics

    huggingface.co1 month agoView details

  1. ggml-org/llama.cpp b10436

    <details open> mtmd, common: various fixes (#27071) * apply fixes * cont * revert gguf fix </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10436/llama-b10436-bin-macos-arm64.tar…

    github.com1 month agoView details

  2. openai/openai-python v3.1.0

    ## [3.1.0](https://github.com/openai/openai-python/compare/v3.0.0...v3.1.0) (2026-08-14) ### Features * **api:** add WebSocket stream IDs ([#3612](https://github.com/openai/openai-python/issues/3612)) ([d9029e3](https://github.com/openai/openai-python/commit/d9029e3ada3c008b4631…

    github.com1 month agoView details

  3. ggml-org/llama.cpp b10435

    <details open> jinja : fix quadratic cost in gather_string_parts (#27034) * jinja : fix quadratic cost in gather_string_parts * fix some comments * remove test </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-or…

    github.com1 month agoView details

  4. langchain-ai/langchain langchain-openrouter==0.2.8

    Changes since langchain-openrouter==0.2.7 release(openrouter): 0.2.8 (#39658) chore(model-profiles): refresh model profile data (#39646) chore(model-profiles): refresh model profile data (#39625) fix(openrouter): preserve cost metadata in usage chunks (#39338) chore(model-profil…

    github.com1 month agoView details

  5. langchain-ai/langchain langchain-core==1.5.5

    Changes since langchain-core==1.5.4 release(core): 1.5.5 (#39655) fix(core): make abatch_iterate consistent with batch_iterate for None and zero size (#39367) fix(core): respect pydantic aliases when validating tool inputs (#39572) fix(core): issues in merging chunks (#39535) fi…

    github.com1 month agoView details

  6. langchain-ai/langchain langchain-openai==1.5.1

    Changes since langchain-openai==1.5.0 release(openai): 1.5.1 (#39653) fix(openai): preserve streamed encrypted reasoning (#39635) chore(infra): support langsmith gateway in CI (#39651)

    github.com1 month agoView details

  7. ggml-org/llama.cpp b10427

    <details open> sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779) Measured on Arc Pro B70 (Battlemage, Level Zero), llama-bench -r 20, two interleaved rounds, tg128: qwen2.5-3B-Instruct Q4_K_M 154.18 -> 158.53 t/s +2.8% gemma-2-2b-it Q4_K_M 162.45 -> 165.62…

    github.com1 month agoView details

  8. ggml-org/llama.cpp b10426

    <details open> ggml: force single thread on wasi (#25686) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10426/llama-b10426-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64…

    github.com1 month agoView details

  9. ggml-org/llama.cpp b10425

    <details open> sycl: fuse the gated-delta-net state writeback cpy (#26643) Port of https://github.com/ggml-org/llama.cpp/pull/23940. Arc Pro B70, Qwen 3.6 27B Q4_K - Medium (48 of its 64 blocks run gated_delta_net), -ngl 99 -fa 1 -ctk f16 -ctv f16 -b 2048 -ub 2048, interleaved A…

    github.com1 month agoView details