Skip to content

Open source / #6680

Open source

ggml-org/llama.cpp b10456

GitHub Releases · github-actions[bot] · 1 month ago

<details open> sycl: fix thread/block count in quantized cpy kernel launches (#27160) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improvement is the q4_0 -> f32 path, with, on a Arc 70, throughput goe…

Source adapter
github-releases
Source weight
3
Points
0
Age
772.4 h
Stored score
0.00001
Current score
0.00001
First seen
2026-08-17T09:03:47.874Z

Newsletter

Get practical AI engineering notes

Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.