Open source / #6680
Open source
ggml-org/llama.cpp b10456
GitHub Releases · github-actions[bot] · 1 month ago
<details open> sycl: fix thread/block count in quantized cpy kernel launches (#27160) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improvement is the q4_0 -> f32 path, with, on a Arc 70, throughput goe…
Source and ranking details
- Source adapter
- github-releases
- Source weight
- 3
- Points
- 0
- Age
- 772.4 h
- Stored score
- 0.00001
- Current score
- 0.00001
- First seen
- 2026-08-17T09:03:47.874Z
Newsletter
Get practical AI engineering notes
Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.