Skip to content

Open source / #22123

Open source

ggml-org/llama.cpp v0.4.0

GitHub Releases · github-actions[bot] · 13 days ago

## Overview llama.cpp 0.4.0 adds initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options, and a ggml update to 0.23.0 with major sparse flash attention and RDMA work. ### API changes - Added `llama_l…

Source adapter
github-releases
Source weight
3
Points
0
Age
326.1 h
Stored score
0.00006
Current score
0.00006
First seen
2026-09-05T08:37:05.904Z

Newsletter

Get practical AI engineering notes

Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.