Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1063 commits
Select commit Hold shift + click to select a range
6dbbac4
opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555)
lhez Sep 29, 2026
cee37ff
ci: add zdnn backend build but not test (#29541)
taronaeo Sep 29, 2026
748d422
ggml-cuda: HIP: optimize packed byte subtraction (`__vsubss4` -> `__v…
thelittlefireman Sep 29, 2026
6a2743f
CUDA: bitonic argsort handles rows wider than one block (#28957)
ServeurpersoCom Sep 29, 2026
7fee178
hexagon: optimize concat op (#29673)
trivikram-reddy1 Sep 29, 2026
48de2a1
model : support classifier_pooling for rerankers (#29627)
boshjerns Sep 29, 2026
d3954b9
ggml : check row bounds in get_rows_back (#29575)
angt Sep 29, 2026
a6ea155
gguf : reject tensor size that wraps after padding (#26979)
x14ngch3n Sep 29, 2026
19e28a2
Hexagon f16 activation ops (#29209)
cqderek Sep 29, 2026
eae11d2
ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
taronaeo Sep 30, 2026
931351e
vendor: update BoringSSL to 0.20260929.0 (#29669)
cabelo Sep 30, 2026
649dcb1
add GLM-5.3-Flash (GLM5-Next) support (#27773)
timkhronos Sep 30, 2026
2a53ace
SYCL: reduce tensor allreduce sync with pinned host buffers (#29604)
Captain-Tripps Sep 30, 2026
72db1e0
ci : add models backend check (#29651)
CISC Sep 30, 2026
272aad8
musa : define __CUDA_ARCH__ for device passes (#29508)
yeahdongcn Sep 30, 2026
db00347
ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
ServeurpersoCom Sep 30, 2026
25747b0
openvino: serve GET_ROWS on a weight view from the base Constant (#28…
ServeurpersoCom Sep 30, 2026
fa2bde5
ui : type-safe API types, fetch helpers and download-ready models sto…
allozaur Sep 30, 2026
f653250
ui : model id grammar for sidecars, quants and capability parsing (#2…
allozaur Sep 30, 2026
9b43336
ui : Hugging Face Hub data layer (#27947)
allozaur Sep 30, 2026
4cfb6d1
ui : model memory-fit estimation (#27957)
allozaur Sep 30, 2026
8664eae
ui : model download pipeline (#27959)
allozaur Sep 30, 2026
4a096b8
ui : shared model display primitives (#29644)
allozaur Sep 30, 2026
8df332d
model-conversion : add --add-bos to run org model script (#29558)
danbev Sep 30, 2026
90c908d
cpu: accept BF16 in src1 of mul_mat (#28937)
ServeurpersoCom Sep 30, 2026
2090f60
ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
am17an Sep 30, 2026
185103d
llama: llama_prefetch_rows (#29599)
am17an Sep 30, 2026
3b3d022
ci : fix Models Backend Check by shortening the hrm_text fixture (#29…
ServeurpersoCom Sep 30, 2026
bdeb855
ggml-et : remove useless alloca() (#29663)
angt Sep 30, 2026
ca2e203
jinja : support coerced array attributes (#29574)
CISC Sep 30, 2026
22bdcc4
mimo : support dflash (convert + feature extraction) (#29650)
ggerganov Sep 30, 2026
b046420
cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#2…
ServeurpersoCom Sep 30, 2026
876c75b
codeowners : remove former ZenDNN owner (#29747)
vishalMCE Sep 30, 2026
2149c00
ggml/gguf : fix integer overflow (#29384)
apach301 Sep 30, 2026
05af0d2
glm5-next: give dead indexer slots unique scatter rows (#29745)
ServeurpersoCom Sep 30, 2026
60e9cf7
batch: migrate the rest of examples to llama_batch_ext (#29601)
ngxson Sep 30, 2026
81ff93e
llama: properly handle KV on training (#28520)
ngxson Sep 30, 2026
b016f46
convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
Swigler Sep 30, 2026
4f31296
test-llama-archs : toggle causal_attn to catch graph shape changes (#…
sihanyu03 Sep 30, 2026
4453b53
llama : preserve original batch order for speculative decoding layer …
hthadicherla Sep 30, 2026
feb9a3d
args: fix cli download mmproj arg (#28977)
pr3pony Sep 30, 2026
a4d880f
Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
ebateni Sep 30, 2026
f872b59
cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
yeahdongcn Sep 30, 2026
f7b384c
ggml-opencl : replace alloca() with std::vector (#29765)
angt Sep 30, 2026
0c1e570
webgpu: fix SSM_SCAN binding aliasing (#29750)
yomaytk Oct 1, 2026
10f340d
model : re-enable -sm tensor for qwen4exp (#28569)
kh0pper Oct 1, 2026
66bcc27
docs : refresh CPU ops support matrix (#29666)
CaramelizedCUDA Oct 1, 2026
79625e0
llama-bench : fix docs (#29464)
mairp Oct 1, 2026
2232bc8
metal : use bf16 math for mxfp4 mul-mat (#29770)
ggerganov Oct 1, 2026
7dad6db
llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
kushalgarg101 Oct 1, 2026
db33d3c
vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
tokinasin Oct 1, 2026
3ec4df4
opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler …
lhez Oct 1, 2026
b8f96c3
common : add LLM-jp-4.1 Harmony dialect handler (#29681)
e-mon Oct 1, 2026
b0aca3c
BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
pradeeptrgit Oct 1, 2026
3aa0ce9
hex-workqueue: fix race condition in seqn getting out of sync with id…
max-krasnyansky Oct 1, 2026
f11d642
HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma d…
IMbackK Oct 1, 2026
32dd62e
llama-mmap : avoid a second full-size copy of each tensor with direct…
praneshgo Oct 1, 2026
def4d40
jinja : skip copying loop scope unless a loop filter needs it (#29776)
xarillian Oct 1, 2026
5503b04
meta: clear inactive AllReduce shards with FILL, not SCALE (#29793)
ggerganov Oct 1, 2026
552f18f
mtmd: cap max_image to ubatch for non_causal models (#29773)
ngxson Oct 1, 2026
7677678
CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
ORippler Oct 1, 2026
66e0c17
llama: fix qwen4exp (#29751)
am17an Oct 1, 2026
c061df1
Qwen4Exp: add MTP (#29761)
am17an Oct 1, 2026
b56f34a
CUDA: Handle compute type for NVFP4 on cublass path (#29173)
ynankani Oct 1, 2026
869034b
llama : fix invalid assert in recurrent memory (#29799)
ggerganov Oct 1, 2026
4b1622a
webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 …
yomaytk Oct 1, 2026
13b4d71
metal : release temporary private transfer buffers (#29777)
mvanlamz Oct 1, 2026
42d9581
cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
tkittich Oct 1, 2026
2b36825
convert : write Gemma embedding scale for DFlash drafts (#29802)
kabu1204 Oct 1, 2026
d775ebf
server: return HTTP 400 for invalid embedding requests (#29060)
SamMalayek Oct 1, 2026
dcd387a
hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT v…
njsyw1997 Oct 1, 2026
81e39ad
llama : clamp kpool re-pool bound to existing pools (#29805)
ggerganov Oct 1, 2026
e358d59
ci: fix Fusion / metal by updating the qwen4exp baseline (#29812)
ServeurpersoCom Oct 1, 2026
68e79bd
skill: note about model-specific CLI arguments + testings (#29808)
ngxson Oct 1, 2026
f1cee99
common,rpc : fix cache dir creation through symlinks on buggy libstdc…
angt Oct 1, 2026
78e2964
llama: refer to segment documentation [no ci] (#29074)
JohannesGaessler Oct 1, 2026
ec7630a
CUDA: fix 2 broken Volta FA cases (#29803)
JohannesGaessler Oct 1, 2026
a868c3e
hexagon: add q2_k and q3_k quant type support (#29717)
jhen0409 Oct 1, 2026
159c651
qwen4exp: fix tests (#29819)
am17an Oct 2, 2026
5fc4f3c
hexagon: install rebuilt HTP skels (#29828)
kurquhar Oct 2, 2026
207bdab
pyproject : add linux platform marker to uv torch source (#29177)
Yezat Oct 2, 2026
fb4b273
vulkan: add logging to pipeline compile issues (#29794)
0cc4m Oct 2, 2026
254b177
ci : fix missing zdnn backend check (#29837)
taronaeo Oct 2, 2026
631109b
ggml : add `alloc_buffer_n` to buffer type interface (#23671)
ggerganov Oct 2, 2026
4e2713c
qwen4exp : optimize mask constructions (#29824)
ggerganov Oct 2, 2026
c328acc
sycl : do not use slow oneDNN reference matmul and fattn (#28985)
lslusarczyk Oct 2, 2026
b933289
sycl: large register file for D=512 FA vec kernels (#29062)
Titaniumtown Oct 2, 2026
9e258a6
vulkan: disable large matmul tile on Samsung GPUs with 32KB shared me…
jwsong98 Oct 2, 2026
392ded6
SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
cwriter Oct 2, 2026
a8c9a4e
opencl: use sigmoid f16 for bf16 (#29787)
lhez Oct 2, 2026
6805ae3
llama : use GGML_ABORT instead of throw (#29840)
angt Oct 2, 2026
8d81559
llama : silence unused-result warnings (#29839)
angt Oct 2, 2026
70849ee
common : remove fs_open_ifstream() by using u8path() (#29841)
angt Oct 2, 2026
a4cb4c6
llama, server: add /v1/systemone API (models: laya, julia-1, lev, ope…
ngxson Oct 2, 2026
926862e
metal : add tensor API flash attention kernel for F16 KV (#29570)
ethanhq Oct 2, 2026
d8fbd25
readme : add cmd install commands (#29850)
ggerganov Oct 2, 2026
46ca246
model: support nimble decision model (#29844)
ngxson Oct 2, 2026
dd4c286
ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
devYRPauli Oct 2, 2026
2923cf2
ggml-quants : avoid invalid rounding in qkx3 scale search (#29817)
devYRPauli Oct 2, 2026
134b2bb
ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#2…
devYRPauli Oct 2, 2026
1fb7ef3
spec : add probabilistic sampling for simple draft and MTP (#27694)
praneshgo Oct 2, 2026
4ebdf2c
ci : use t4-medium for cuda jobs (#29842)
CISC Oct 2, 2026
bed0a85
CUDA: fuse shared experts into MMVQ (#29184)
am17an Oct 2, 2026
99b9548
model: add support for clef decision model (text-only) (#29831)
ngxson Oct 3, 2026
889edf4
qwen4exp : halve the indexer score memory (#29825)
ServeurpersoCom Oct 3, 2026
cb7934c
model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
tdakhran Oct 3, 2026
b92761a
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, …
ravi9 Oct 3, 2026
436f6f8
graph: gather the recurrent states once so the reserve covers every s…
ServeurpersoCom Oct 3, 2026
a55e952
ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904)
ServeurpersoCom Oct 3, 2026
9bf55f4
chat : honor json_schema in Ling 3.0 parser (#29813)
devYRPauli Oct 3, 2026
edd6e2b
common : add common_is_tty() helper and fix deprecated warnings on Wi…
angt Oct 3, 2026
1537a0a
server : fix laya abort by limiting n_batch to n_ubatch (#29903)
notforu777 Oct 3, 2026
eec18f5
vendor : update cpp-httplib to 0.59.0 (#29886)
cabelo Oct 3, 2026
836d571
mtmd : fix deprecated strdup warning on Windows (#29863)
angt Oct 3, 2026
11fe021
webgpu: add f16 support to fill/set_rows (#29897)
yomaytk Oct 4, 2026
f98b31c
ci : improve release flow (#29913)
CISC Oct 4, 2026
0faee50
ci : pushing tag needs deploy key (#29937)
CISC Oct 4, 2026
bf9a0cc
server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#2…
angt Oct 4, 2026
6716df6
common : prepare load_from_models_dir() for path conversion (#29674)
angt Oct 4, 2026
8330e96
spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924)
praneshgo Oct 4, 2026
0504396
imatrix: calculate activation-based statistics for new format (GGUF) …
EAddario Oct 4, 2026
16c163d
vulkan: fix rdna4 mat_vec tuning (#29934)
0cc4m Oct 4, 2026
dd26678
CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941)
JohannesGaessler Oct 4, 2026
2bc5635
cuda : move blocks_per_col to where it is used (#29939)
angt Oct 4, 2026
46847e6
ci : set default permissions (#29945)
CISC Oct 4, 2026
dbe4c3e
chat-peg-parser : clear current_tool when pending_tool_call is reset …
T-Anas Oct 4, 2026
bf79dbb
AGENTS.md : revamp (#29656)
am17an Oct 4, 2026
7f2dd88
ci : add windows arm64 vulkan release (#29954)
CISC Oct 4, 2026
2e7c58c
ci : windows llvm build requires ninja multi-config (#29959)
CISC Oct 4, 2026
0eb6d9a
cuda : move neu_padded to where it is used (#29940)
angt Oct 4, 2026
a7b94df
ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806)
SongXiaoXi Oct 4, 2026
2ca15f5
CUDA: refactor swizzling code (#29612)
JohannesGaessler Oct 4, 2026
0bb496d
llama: support both embd + raw tokens in batch (#29622)
ngxson Oct 4, 2026
a7fb71f
log, server: self contained colors, split child commands from logs in…
ServeurpersoCom Oct 4, 2026
d89651a
CUDA: prefer whole-tile FlashAttention scheduling for efficient two-s…
anujj Oct 5, 2026
9d3aba6
CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633)
ynankani Oct 5, 2026
a3a1c47
metal : few-row MMA mat-mul (#29869)
pratiknarola-t Oct 5, 2026
1b43d31
cuda: tile the lightning indexer over keys and tokens for 4 heads (#2…
ServeurpersoCom Oct 5, 2026
8216c84
webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29…
yomaytk Oct 5, 2026
ebe18be
vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (…
EAddario Oct 5, 2026
9f12cd4
ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)
alanhc Oct 5, 2026
4ca6b76
ci : fix docker workflow permissions (#29979)
CISC Oct 5, 2026
e5983d6
ci : winget urls must be separate strings (#29978)
CISC Oct 5, 2026
2107910
kv-cache: fix restoring mismatched KV cache rotation by saving exact …
eapache Oct 5, 2026
c173a53
llama : fix unexpected graph reallocation in the k-pool models (#29958)
ggerganov Oct 5, 2026
b3daa07
vulkan: sparse flash attention for quantized K/V (#29639)
fxgsell Oct 5, 2026
806eee9
vulkan: fix stale prealloc_y reuse across flash attention and soft_ma…
fxgsell Oct 5, 2026
8e16421
server: reject partial media truncation (#24076)
he-yufeng Oct 5, 2026
2ed93db
ci : disable unused qemu in docker build (#29984)
CISC Oct 5, 2026
8b2fbaf
CUDA: Optimize accumulation in mmq for NVFP4 type (#29857)
kmorennv Oct 5, 2026
9871df5
server: support vision input for Clef (#29969)
ngxson Oct 5, 2026
8f9ae20
ci : disable failing test on virtual Metal device (#29993)
ggerganov Oct 5, 2026
9d853bb
webui: Use toLocaleString() format consistently across chat message s…
grivera64 Oct 5, 2026
994e8f2
ci : add "Require Docker" flag to make-release workflow (#29989)
ggerganov Oct 5, 2026
b809b88
cuda: use the vector lightning indexer kernel on MUSA (#29990)
ServeurpersoCom Oct 5, 2026
3c9e747
vulkan: revert mul_mat_id tile selection PR #29182 (#29936)
virajwad Oct 5, 2026
6c59c40
vulkan: fix Flash Attention shmem write out of bounds (#29988)
0cc4m Oct 5, 2026
e117148
CUDA: make the alloc_deps check batch independent (#29986)
am17an Oct 5, 2026
f05c8b2
ggml : bump version to 0.26.0 (ggml/1652)
ggerganov Oct 5, 2026
c06f841
sync : ggml
ggerganov Oct 5, 2026
4d60b4d
common, server : report model input/output modalities in GET /models …
angt Oct 5, 2026
d812350
llama.cpp : bump version to 0.6.0 (#29997)
ggerganov Oct 5, 2026
8345f33
hexagon: matmul and flash-atten scalability updates (#29974)
max-krasnyansky Oct 5, 2026
c250304
ci : skip container re-tagging when Require Docker is disabled (#30008)
ggerganov Oct 5, 2026
7049ff0
ci : add 1accel label [no ci] (#30016)
CISC Oct 5, 2026
50569eb
hexagon: add pool op support (#29995)
aparmp-quic Oct 5, 2026
5e03bdd
hexagon: ssm-conv updates (#29971)
tboinovski1 Oct 6, 2026
43fe9c6
llama: fix k-pool scatter data race on shared sequences (#29994)
ServeurpersoCom Oct 6, 2026
b9a5a00
ggml-openvino: fix CI tests; fix GPU regressions. (#30037)
ravi9 Oct 6, 2026
63bef27
vendor : update LibreSSL to 4.3.3 [no ci] (#30019)
cabelo Oct 6, 2026
cbb7d52
test-llama-archs : initialize backends before generating models (#30034)
booxter Oct 6, 2026
6753a03
ggml: refactor selective expert copying to user code (#29943)
am17an Oct 6, 2026
1a3011c
llama : re-reserve the sched when the nextn extraction flags change (…
pwilkin Oct 6, 2026
6c73b3e
convert : add text_config as fallback [transformers 5.18] (#30040)
frozenblade1224 Oct 6, 2026
d7a695e
scripts : limit apiabi checks to libllama and libmtmd (#30038)
danbev Oct 6, 2026
f0c41e0
models : consolidate nextn row cropping into shared helpers (#30017)
ggerganov Oct 6, 2026
4f54067
HIP: use -O0 for host code in debug builds (#29795)
JohannesGaessler Oct 6, 2026
58cb913
vulkan : check for null vkEnumerateInstanceVersion (#29872)
ianloic Oct 6, 2026
a043d38
metal : fix excess threadgroup memory in quantized flash attention (#…
masterFoad Oct 6, 2026
da263e7
models: support pplx-decider (#30044)
ngxson Oct 6, 2026
ab09ea4
cuda: BF16/FP16 conversion to f32 chunking (#29442)
thelittlefireman Oct 6, 2026
65840ed
ggml: fix CLAMP on non-contiguous views (CPU, CUDA) (#29517)
ServeurpersoCom Oct 6, 2026
a46709b
RPC: add `-sm tensor` (#26610)
am17an Oct 6, 2026
2207c8e
opencl: fix OOB read in adreno xmem GEMM (#30041)
lhez Oct 6, 2026
4fbc76d
model: support embeddinggemma2 (text+vision+audio) (#30054)
ngxson Oct 6, 2026
3109914
llama: remove the gather path of the glm5-next sparse attention (#30042)
ServeurpersoCom Oct 6, 2026
4625240
model : add K2 Horizon dense and MoVA support (#29535)
bitalov Oct 6, 2026
abeada3
vocab : implement PLaMo-3 tokenizer pre-segmentation (#30045)
tokinasin Oct 6, 2026
51ce9c1
ggml-cuda: use per-thread stream for buffer-init padding memset (#28782)
harkgill-amd Oct 6, 2026
5ad1c5d
cuda : add BF16 support for XIELU (#29955)
Qiao12-pixel Oct 6, 2026
c479922
hexagon: CPY/CONCAT/CONT/DUP overhaul to use DMA/HVX for all cases (#…
max-krasnyansky Oct 6, 2026
f498f86
ggml-webgpu: no dawn native features on wasi (#27069)
MendyBerger Oct 7, 2026
78651c4
sycl: add IQ3_S multi-column MMVQ (#29500)
clemenswasser Oct 7, 2026
4d756bc
vulkan: fix amd iGPU slow checkpoint read (#30049)
0cc4m Oct 7, 2026
5e5b628
sycl : fattn_kv_buffers cleanup (#27689)
Pratyush-gg Oct 7, 2026
d2a79e6
sycl: accelerate GLM MLA prefill with MKL flash attention (#29171)
anantshri Oct 7, 2026
005a1e1
[SYCL] fix the issue in mixed different model GPUs in FA (#29071)
arthw Oct 7, 2026
fa3c2fa
tests : retain the anchor when testing recurrent rollback (#29923)
eapache Oct 7, 2026
36a7391
ggml-webgpu: fix flash_attn supports_op check for overlapping KV (#28…
yomaytk Oct 7, 2026
2690873
imatrix : include clocale for std::setlocale (#30079)
dazzywi Oct 7, 2026
b7dafa0
vendor : update cpp-httplib to 0.60.0 (#30081)
cabelo Oct 7, 2026
ad21565
musa: use the tile lightning indexer kernel (#30080)
yeahdongcn Oct 7, 2026
7481354
convert : fix token configuration for PLaMo-3 (#29843)
tokinasin Oct 7, 2026
42b021b
vocab : add plamo fim tokens (#30090)
CISC Oct 7, 2026
d0b490f
sampling : use greedy selection for eligible temperature-zero chains …
hthadicherla Oct 7, 2026
48499d2
qwen3tts : guard speaker_encoder_config patch for CustomVoice variant…
SIDDARTHAREDDY8 Oct 7, 2026
b9acf13
feat: add GLM5Next MTP, optimize (#29928)
pwilkin Oct 7, 2026
7e8324f
metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT …
nyo16 Oct 7, 2026
9881906
metal : few-row MMA mat-mul for the remaining src0 types (#30065)
pratiknarola-t Oct 7, 2026
448147d
llama: share the nextn tensor flags between models (#30097)
ServeurpersoCom Oct 7, 2026
18b5f8b
server : accumulate generated text and tokens as parse input (#29876)
aldehir Oct 7, 2026
42c787e
cuda: update uncoalesced memory reads in pool2d (#29425)
shenron0101 Oct 7, 2026
d6cf9ac
llama : add a GPU cache for MoE experts kept in host memory (#29887)
am17an Oct 7, 2026
50a6c5c
mtmd: add cohere2 vision support (#30062)
Terrencezzj Oct 7, 2026
b86d2f0
cuda: FWHT kernels for block widths above 512 (#29100)
bri-prism Oct 7, 2026
88dcc46
model : add LiquidAI/d1-3B decision model (#30110)
tdakhran Oct 7, 2026
5de7334
chat : name tool and argument parser rules by index (#30088)
Frost-54 Oct 7, 2026
bd4eeaa
chat: fix jinja parser for TranslateGemma (#30096)
nbud Oct 7, 2026
7081510
hexagon: improved GELU accuracy (#30104)
kurquhar Oct 7, 2026
aa5e009
hexagon: support tiled Q4_K and Q6_K GET_ROWS (#30115)
kurquhar Oct 8, 2026
a657f7e
model : add LiquidAI/d1-omni-600M decision model (#30114)
tdakhran Oct 8, 2026
06cad0b
hexagon: Q6_K weight dequant speedup (#30121)
kurquhar Oct 8, 2026
9c2e0e4
hexagon: enable alloc_buffer_n (#30126)
max-krasnyansky Oct 8, 2026
9b4ed0c
sycl: remove duplicate block-size defines from op headers (#29507)
Titaniumtown Oct 8, 2026
75118a3
convert : support Qwen3.5 embedding models (#27920)
SamMalayek Oct 8, 2026
847f447
sycl: add grouped MoE XMX GEMM (#29245)
cwriter Oct 8, 2026
24e4183
ggml-cuda: assign four GDN state columns per warp (#30087)
SongXiaoXi Oct 8, 2026
37ac634
model : support classifier_activation for rerankers (#29692)
boshjerns Oct 8, 2026
000bee5
sycl: stage bulk uploads (model loading) through a pinned ring buffer…
cwriter Oct 8, 2026
ff30363
sycl: fuse the delta-net alpha gate (add + unary + mul) (#29687)
Titaniumtown Oct 8, 2026
fda1866
CUDA: fix norm family kernels when ne[2]/ne[3] exceed grid dim limits…
edenfunf Oct 8, 2026
d888016
cuda : Use byte strides for roll to allow non-contiguous ROLL operati…
bertaye Oct 8, 2026
097f5b5
hex-mmadd: do not assume aligned read/write when bias-add is fused (#…
max-krasnyansky Oct 8, 2026
46baf1f
sycl: FWHT optimizations (#29605)
bri-prism Oct 8, 2026
08246a2
cuda : support arbitrary striding for unary ops on f16, f32, and bf16…
saady789 Oct 8, 2026
03aa006
vendor : update cpp-httplib to 0.60.1 (#30134)
angt Oct 8, 2026
dac3087
vulkan: extend sparse FA support to coopmat2 (#30003)
jeffbolznv Oct 8, 2026
ff5888f
vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values (…
gianni-cor Oct 8, 2026
033df86
server : preserve context checkpoints across slot save/restore (#26004)
Tough-Respawn Oct 8, 2026
c811cb8
llama: support MoE cache over multiple GPUs (#30112)
am17an Oct 8, 2026
4f92965
ui: apply ui_settings on first visit in router mode (#29668)
simonether Oct 8, 2026
1167d3f
CUDA: fix CCCL version guard breaking on major version rollover (#29453)
hey-gm Oct 8, 2026
c35b667
CUDA : looped PAD kernel for more than 65535 rows or slices (#30147)
pskrunner14 Oct 8, 2026
fc9ce6b
CUDA: fix MMQ out-of-bounds reads (#29953)
JohannesGaessler Oct 8, 2026
a11f57b
model : fix DFlash output head sharing (#30111)
hthadicherla Oct 8, 2026
71ad059
CUDA: improve top-k algorithm selection (#28713)
praneshgo Oct 8, 2026
de7fa0a
Musa FWHT fix (#30167)
bri-prism Oct 8, 2026
14fd0bf
merge: sync upstream llama.cpp onto the b10453.6 decision line
Siddhesh2377 Oct 8, 2026
120bb61
tools: remove clef-demo
Siddhesh2377 Oct 8, 2026
c6aa47c
llama: score GLiNER Decide on the DeBERTa-v3 encoder
Siddhesh2377 Oct 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
6 changes: 3 additions & 3 deletions .devops/intel.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
ARG ONEAPI_VERSION=2025.3.3-0-devel-ubuntu24.04
ARG ONEAPI_VERSION=2026.1.1-devel-ubuntu24.04
ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
ARG APP_REVISION=N/A
Expand All @@ -19,7 +19,7 @@ RUN npm ci
COPY tools/ui/ ./
RUN LLAMA_BUILD_NUMBER="$APP_VERSION" npm run build

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS build
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS build

ARG GGML_SYCL_F16=ON
ARG LEVEL_ZERO_VERSION=1.28.2
Expand Down Expand Up @@ -59,7 +59,7 @@ RUN mkdir -p /app/full \
&& cp requirements.txt /app/full \
&& cp .devops/tools.sh /app/full/tools.sh

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS base
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS base

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down
15 changes: 10 additions & 5 deletions .devops/musa.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,10 +1,9 @@
ARG UBUNTU_VERSION=22.04
# This needs to generally match the container host's environment.
ARG MUSA_VERSION=rc4.3.0
# Target the MUSA build image
ARG BASE_MUSA_DEV_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-devel-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_DEV_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-devel-ubuntu${UBUNTU_VERSION}-s5000

ARG BASE_MUSA_RUN_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_RUN_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-runtime-ubuntu${UBUNTU_VERSION}-s5000

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down Expand Up @@ -37,7 +36,10 @@ RUN apt-get update && \
python3-pip \
git \
libssl-dev \
libgomp1
libgomp1 \
musa-mualg-5-2 \
musa-muthrust-5-2 \
libmthreads-compute

WORKDIR /app

Expand Down Expand Up @@ -80,13 +82,16 @@ LABEL org.opencontainers.image.created=$BUILD_DATE \
org.opencontainers.image.source=$IMAGE_SOURCE

RUN apt-get update \
&& apt-get install -y libgomp1 curl ffmpeg \
&& apt-get install -y libgomp1 curl ffmpeg libmthreads-compute \
&& apt autoremove -y \
&& apt clean -y \
&& rm -rf /tmp/* /var/tmp/* \
&& find /var/cache/apt/archives /var/lib/apt/lists -not -name lock -type f -delete \
&& find /var/cache -type f -delete

# The MUSA runtime image does not register its library directory
RUN echo "/usr/local/musa/lib" > /etc/ld.so.conf.d/musa-runtime.conf && ldconfig

COPY --from=build /app/lib/ /app

### Full
Expand Down
12 changes: 6 additions & 6 deletions .devops/nix/package.nix
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@
]
&& blas.meta.available,
useCuda ? config.cudaSupport,
useMetalKit ? stdenv.isAarch64 && stdenv.isDarwin,
useMetalKit ? stdenv.hostPlatform.isAarch64 && stdenv.hostPlatform.isDarwin,
# Increases the runtime closure size by ~700M
useMpi ? false,
useRocm ? config.rocmSupport,
Expand Down Expand Up @@ -92,7 +92,7 @@ let

cudaBuildInputs = with cudaPackages; [
cuda_cudart
cuda_cccl # <nv/target>
cccl # <nv/target>
libcublas
];

Expand Down Expand Up @@ -166,7 +166,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {
# `xcrun` is used find the path of the Metal compiler, which is varible
# and not on $PATH
# see https://github.com/ggml-org/llama.cpp/pull/6118 for discussion
__noChroot = effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders;
__noChroot = effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders;

nativeBuildInputs =
[
Expand All @@ -181,10 +181,10 @@ effectiveStdenv.mkDerivation (finalAttrs: {
autoAddDriverRunpath
]
++ optionals (effectiveStdenv.hostPlatform.isGnu && enableStatic) [ glibc.static ]
++ optionals (effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];
++ optionals (effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];

buildInputs =
optionals effectiveStdenv.isDarwin darwinBuildInputs
optionals effectiveStdenv.hostPlatform.isDarwin darwinBuildInputs
++ optionals useCuda cudaBuildInputs
++ optionals useMpi [ mpi ]
++ optionals useRocm rocmBuildInputs
Expand Down Expand Up @@ -245,7 +245,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {

# Configurations that are known to result in build failures. Can be
# overridden by importing Nixpkgs with `allowBroken = true`.
broken = (useMetalKit && !effectiveStdenv.isDarwin);
broken = (useMetalKit && !effectiveStdenv.hostPlatform.isDarwin);

description = "Inference of LLaMA model in pure C/C++${descriptionSuffix}";
homepage = "https://github.com/ggml-org/llama.cpp/";
Expand Down
23 changes: 13 additions & 10 deletions .devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
ARG OPENVINO_VERSION_MAJOR=2026.2.1
ARG OPENVINO_VERSION_FULL=2026.2.1.21919.ede283a88e3
ARG OPENVINO_VERSION_MAJOR=2026.4.1
ARG OPENVINO_VERSION_FULL=2026.4.1.22982.07f9c262b05
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.36.3
ARG IGC_VERSION_FULL=2_2.36.3+21719
ARG COMPUTE_RUNTIME_VERSION=26.22.38646.4
ARG COMPUTE_RUNTIME_VERSION_FULL=26.22.38646.4-0
ARG IGC_VERSION=v2.41.5
ARG IGC_VERSION_FULL=2_2.41.5+22716
ARG COMPUTE_RUNTIME_VERSION=26.35.39758.10
ARG COMPUTE_RUNTIME_VERSION_FULL=26.35.39758.10-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
ARG NPU_DRIVER_VERSION=v1.33.0
ARG NPU_DRIVER_FULL=v1.33.0.20260529-26625960453
ARG LIBZE1_VERSION=1.27.0-1~24.04~ppa2
ARG NPU_DRIVER_VERSION=v1.38.0
ARG NPU_DRIVER_FULL=v1.38.0.20260910-34487311128
ARG LIBZE1_VERSION=1.32.0-1~24.04~ppa1

# Optional proxy build arguments
ARG http_proxy=
Expand Down Expand Up @@ -90,6 +90,9 @@ RUN bash -c "source ${OpenVINO_DIR}/setupvars.sh && \
cmake -B build/ReleaseOV -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DLLAMA_BUILD_TESTS=OFF \
-DGGML_NATIVE=OFF \
-DGGML_BACKEND_DL=ON \
-DGGML_CPU_ALL_VARIANTS=ON \
-DGGML_OPENVINO=ON && \
cmake --build build/ReleaseOV --parallel "

Expand Down Expand Up @@ -170,7 +173,7 @@ RUN --mount=type=cache,target=/var/cache/intel-npu,sharing=locked \
fi; \
DEB=/var/cache/intel-npu/libze1_${LIBZE1_VERSION}_amd64.deb; \
if [ ! -f "$DEB" ]; then \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260324T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260830T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
fi; \
mkdir /tmp/npu/ && cd /tmp/npu/ && tar -xf "$TGZ" && cp "$DEB" .; \
apt-get update; \
Expand Down
2 changes: 1 addition & 1 deletion .ecrc
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"Exclude": ["^\\.gitmodules$", "stb_image\\.h"],
"Exclude": ["^\\.gitmodules$", "stb_image\\.h", "examples/test-cmake/build/", "examples/test-cmake/build-subdir/"],
"Disable": {
"IndentSize": true
}
Expand Down
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
blank_issues_enabled: true
blank_issues_enabled: false
contact_links:
- name: Got an idea?
url: https://github.com/ggml-org/llama.cpp/discussions/categories/ideas
Expand Down
95 changes: 95 additions & 0 deletions .github/actions/ccache-buckets/action.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
name: "ccache-buckets"
description: "Save/restore latest GitHub Actions ccache matching a key prefix to/from HF buckets"
inputs:
key:
description: "Cache key prefix to match and load"
required: true
folder:
description: "Bucket folder containing ccache files"
required: true
evict-old-files:
description: "Corresponds to the ccache --evict-older-than AGE option, where AGE is the number of seconds or days followed by the 's' or 'd' suffix respectively."
default: ''
save:
description: "Save ccache"
required: false
default: false
type: boolean
hf_bucket:
description: 'Hugging Face buckets path'
required: true

runs:
using: "composite"
steps:
- name: Install Hugging Face Hub CLI
shell: bash
run: |
python3 -m venv .venv-hf
.venv-hf/bin/pip install -U huggingface_hub==1.28.0
- name: Restore ccache from buckets
if: ${{ inputs.save != 'true' }}
shell: bash
run: |
set +e -uo pipefail
source .venv-hf/bin/activate
CCACHE_DIR=$(ccache -k cache_dir)
if [[ -d "$CCACHE_DIR" ]]; then
CACHE_PATH=$(hf buckets list "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}" --json | jq -r '[.[] | select(.type == "file") | select(.path | startswith("${{ inputs.folder }}/${{ inputs.key }}") and endswith(".tar.gz"))] | sort_by(.path) | last | .path // ""')
if [[ -n "$CACHE_PATH" ]]; then
echo "Restoring ccache from '$CACHE_PATH'."
hf buckets cp "hf://buckets/${{ inputs.hf_bucket }}/$CACHE_PATH" ccache_bucket.tar.gz
mkdir -p ccache_bucket
if tar -xzf ccache_bucket.tar.gz -C ccache_bucket; then
rm -rf "$CCACHE_DIR"
mv ccache_bucket "$CCACHE_DIR"
ccache -z
fi
rm ccache_bucket.tar.gz
else
echo "No ccache found."
fi
else
echo "'$CCACHE_DIR' not found."
fi
- name: Save ccache to buckets
if: ${{ inputs.save == 'true' }}
shell: bash
run: |
if [[ -n "$HF_TOKEN" ]]; then
set +e -uo pipefail
source .venv-hf/bin/activate
CCACHE_DIR=$(ccache -k cache_dir)
if [[ -d "$CCACHE_DIR" ]]; then
ccache -s
if [[ -n "${{ inputs.evict-old-files }}" ]]; then
ccache --evict-older-than "${{ inputs.evict-old-files }}"
fi
DATESTAMP=$(date -u +'%Y-%m-%dT%H:%M:%SZ')
CACHEFILE="${{ inputs.key }}-$DATESTAMP.tar.gz"
if tar -czf ccache_bucket.tar.gz -C "$CCACHE_DIR" .; then
hf buckets cp ccache_bucket.tar.gz "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}/$CACHEFILE"
fi
rm ccache_bucket.tar.gz
else
echo "'$CCACHE_DIR' not found."
fi
fi
- name: Remove old ccache files from buckets
if: ${{ inputs.save == 'true' }}
shell: bash
run: |
if [[ -n "$HF_TOKEN" ]]; then
set +e -uo pipefail
source .venv-hf/bin/activate
CACHE_FILES=$(hf buckets list "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}" --json | jq -r '[.[] | select(.type == "file") | select((.uploaded_at | .[:19]+"Z" | fromdateiso8601) < (now - 5 * 60)) | select(.path | startswith("${{ inputs.folder }}/${{ inputs.key }}") and endswith(".tar.gz"))] | sort_by(.path)[:-1] | .[] | [.path // ""] | @tsv')
if [[ -n "$CACHE_FILES" ]]; then
echo "Removing old ccache files..."
while IFS=$'\t' read -r CACHE_PATH; do
hf buckets rm "hf://buckets/${{ inputs.hf_bucket }}/$CACHE_PATH" -y
done <<< "$CACHE_FILES"
fi
fi
48 changes: 38 additions & 10 deletions .github/actions/ccache-clear/action.yml
Original file line number Diff line number Diff line change
@@ -1,22 +1,50 @@
# note: place this as the last step of the job, so the new cache is saved by "Post ccache" right after the old one is cleared
name: "ccache-clear"
description: "Delete all GitHub Actions caches matching a key prefix"
description: "Delete GitHub Actions caches matching a key prefix, oldest first"
inputs:
key:
description: "Cache key prefix to match and delete"
required: true
older:
description: "Only delete caches created more than this long ago (e.g. 90m, 1h, 1d). By default all matching caches are deleted"
required: false
default: ""
min:
description: "Stop deleting if fewer than this many caches would remain (e.g. 1). By default there is no minimum"
required: false
default: "0"
dry-run:
description: "Only print the caches that would be deleted, without deleting them"
required: false
default: "false"

runs:
using: "composite"
steps:
- name: Clear caches
- name: Install GitHub CLI if missing
shell: bash
run: |
CACHES=$(gh cache list --key "ccache-${{ inputs.key }}" --json id,key --jq '.[] | "\(.id) \(.key)"' 2>/dev/null)
if [ -z "$CACHES" ]; then
echo "No caches found with key prefix: ${{ inputs.key }}"
exit 0
# e.g. in container jobs, where it is not preinstalled
if ! command -v gh >/dev/null 2>&1; then
echo "GitHub CLI not found, installing..."
if ! command -v curl >/dev/null 2>&1; then
apt-get update >/dev/null 2>&1 || true
apt-get install -y curl >/dev/null 2>&1 || true
fi
mkdir -p -m 755 /etc/apt/keyrings
curl -fsSL https://cli.github.com/packages/githubcli-archive-keyring.gpg | tee /etc/apt/keyrings/githubcli-archive-keyring.gpg >/dev/null
chmod go+r /etc/apt/keyrings/githubcli-archive-keyring.gpg
echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" > /etc/apt/sources.list.d/github-cli.list
apt-get update >/dev/null 2>&1 || true
apt-get install -y gh || { echo "Failed to install GitHub CLI (gh)" >&2; exit 1; }
fi
while read -r id key; do
echo "Deleting cache: $id ($key)"
gh cache delete "$id"
done <<< "$CACHES"
command -v gh >/dev/null 2>&1 || { echo "GitHub CLI (gh) is required but could not be installed" >&2; exit 1; }

- name: Clear caches
shell: bash
run: |
bash scripts/ccache-clear.sh \
--key "${{ inputs.key }}" \
--older "${{ inputs.older }}" \
--min "${{ inputs.min }}" \
${{ inputs.dry-run == 'true' && '--dry-run' || '' }}
2 changes: 1 addition & 1 deletion .github/actions/get-tag-name/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ runs:
run: |
BUILD_NUMBER="$(git rev-list --count HEAD)"
SHORT_HASH="$(git rev-parse --short=7 HEAD)"
if [[ "${{ env.BRANCH_NAME }}" == "master" ]]; then
if [[ "${{ env.BRANCH_NAME }}" == "master" || "${{ env.BRANCH_NAME }}" == "b${BUILD_NUMBER}" ]]; then
echo "name=b${BUILD_NUMBER}" >> $GITHUB_OUTPUT
else
SAFE_NAME=$(echo "${{ env.BRANCH_NAME }}" | tr '/' '-')
Expand Down
20 changes: 0 additions & 20 deletions .github/actions/linux-setup-vulkan/action.yml

This file was deleted.

Loading
Loading