Repository navigation
Sync upstream llama.cpp onto the b10453.6 decision line - #3
Open
Siddhesh2377 wants to merge 1063 commits into
Open
Siddhesh2377 wants to merge 1063 commits into
Siddhesh2377 wants to merge 1063 commits into
Conversation
* ci: add zdnn backend build but not test Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ci: attempt to run a ubuntu 26.04 container Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ci: set shell to bash Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ci: clean up comments Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ggml-zdnn: fix compiler errors Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * vendor: attempt to ignore warnings from vendored files Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> --------- Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
…sub4`) (ggml-org#29478) * ggml-cuda: HIP: optimize non-saturating packed byte subtraction (`__vsubss4`) * CI: ignore 1 spilled vgpr in fattn_vec --------- Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread per padded column, so any row above 1024 entries launched an invalid block configuration. Each thread now owns several columns, every stage of the network runs all owned columns before the barrier, and the block is capped at 1024 threads. Shared memory becomes the only bound, which supports_op checks against the device instead of a fixed 1024. Rows up to 1024 run the same work as before. Bit-exact with the CUB path on rows of 2048.
* hex-concat: reduce pkts in gather/transpose hot loop gather directly into dst buffer, use special instruction for gather sync * hex-concat: use fastdiv replace calls to sw divide with fastpath * hex-concat: optimize DMA-HVX pipeline and add transpose helpers
* model : support classifier_pooling for ModernBERT rerankers Assisted-by: Claude Opus 5.5 * model : read classifier pooling type in load_hparams Write classifier.pooling_type from _try_set_pooling_type whenever the config has classifier_pooling, and read it in llama_model_base::load_hparams. ModernBERT falls back to mean when it is unspecified. Assisted-by: Claude Opus 5.5 * conversion : only accept cls and mean for classifier_pooling Assisted-by: Claude Opus 5.5 * model : rename classifier_pooling_type to pooling_type_cls Assisted-by: Claude Opus 5.5
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within (alignment - 1) of SIZE_MAX, which silently bypassed the size overflow guard in gguf_init_from_reader. Reject the tensor before padding when nbytes + (alignment - 1) would overflow. Adds a test-gguf handcrafted case (F32, ne = [4, 2^30-1, 2^30+1, 1]) whose ggml_nbytes = 2^64 - 16 lands in the wrap window. Fails on master, passes with the guard.
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU) Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1 must agree on type), and adds F16 per-thread worker functions in act-ops.c mirroring the existing F32 workers, backed by new HVX f16 kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa, hvx_min_scalar_f16 family). SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device (QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's F16 path is code-complete and builds clean on host + all 4 DSP arch variants (v73/v75/v79/v81), but has no F16 test-case coverage in test-backend-ops and is therefore unverified on-device in this change. * hex-ops: align macros * hex-ops: minor formatting --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
cont: fix code style Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx * Add initial MTP support * Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements. * Review driven changes, remove env vars, protect tensors * Strip MTP for initial PR * Clean up after mtp strip * Clean up after mtp strip * Update speculative.cpp * Update llama-context.h * Clean up after mtp strip * Fix tokenizer ignore merges * Improve quantization protection selection * Refactor mhc helpers, graph base * Lint Fixes * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Skip glm5-next in model saver, fix CRLF * Skip glm5-next in sweep * Remove T4 fallback * Review cleanup * Review suggestions * Defer separate MTP gguf handling to MTP PR, drop filter * Repad n_head_kv * kpool init apply * Order by descending score * Drop guard * read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr) * Add glm5-next support to model saver and add arch test fixture * Review cleanup * Kpool pooled caching clarify * Add multi stream support * Finish Rebase * Sparse FA fir DSA prefill * Const * Update llama-model.cpp to fix rebase error * gguf-py : merge tensor map entries for HC tensors * model : use build_gdn_l2_norm in GLM5_NEXT implementation * chore : remove trailing whitespace * model : use new OP precision setting API in GLM5_NEXT implementation * mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit The two clamps around swiglu_split are what ggml_swiglu_clamp already does, so the clamp bounds collapse back to one value. GLM5V also never called set_limit_image_tokens(), so --image-max-tokens had no effect. Assisted-by: Claude Opus 5 (cherry picked from commit 46d18e1) * llama : keep the GLM5-Next k-pool layout across ubatches The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed by the positions relative to the sequence's first one, so the layout now lives on the memory and a ubatch only appends to it. A sequence edit no longer stales every pooled key either, only the ones at or after the edited position, which makes a tail seq_rm free. The pooling subgraph is built unconditionally so the graph shape no longer changes every kpool tokens, and the pool axis is folded into rows before soft_max, which otherwise exceeds the CUDA gridDim.y limit past n_kv 262144. Assisted-by: Claude Opus 5 (cherry picked from commit 5d1c40b) * model : write the GLM5-Next recurrent rollback checkpoints The conv state and the delta net state were only written to the live row, so a rollback restored whatever the checkpoint rows happened to hold. Take the same route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv groups. That also drops a state view that assumed contiguous rows. Enroll the arch in test-recurrent-state-rollback, which catches this under its garbage-filled cache pass. Assisted-by: Claude Opus 5 (cherry picked from commit 5ace37e) * llama: fix PR ggml-org#27773 test-save-load-state restore failure Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence. Assisted-by: Codex * llama: fix PR ggml-org#27773 gpu-rocm graph reallocation Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error. Assisted-by: Codex * llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown Two defects in the cross-ubatch k-pool layout added by the k-pool commit: 1. Wrong results. An edited sequence only rebuilt its pool layout when its cell count changed, so if the first ubatch after an edit added back exactly as many cells as were removed, the stale position-to-cell list survived. With a unified cache and more than one sequence, where another sequence takes the freed cells, the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN). Rebuild whenever the sequence is stale, not only on a size mismatch. 2. Slowdown. "shared" mode was assumed to end only with an edit that forces a rebuild, but sharing also ends when the other sequence is removed. The survivor kept shared = true, pinning cache_safe off and re-pooling every pool on every ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence so one rebuild re-derives sharing and cache_safe returns to 1. Assisted-by: Claude Opus 5 * llama : fix build_attn_mha stream stride for non-contiguous q build_attn_mha split the batch into streams with a stream stride of q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2], when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the stride was then n_head times too large and every stream s >= 1 read another head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was wrong for every stream past the first. Unified KV and decode were unaffected (n_stream == 1, and decode takes the gather path). Other MLA models concat rope so q is contiguous and the computed value is unchanged for them. Compute the stride from the token dimension, which is identical for a contiguous q. Assisted-by: Claude Opus 5 * llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop The shared-cell teardown added to seq_rm (stale every sequence when the layout has shared cells, so a survivor does not keep shared = true and pin cache_safe off) was missing from the other paths that can free shared cells: state_read and state_drop staled only the one sequence. Apply the same re-derivation there and correct the comment that claimed sharing ends only via an edit or seq_rm. Assisted-by: Claude Opus 5 * quant : drop duplicate GLM5-Next hc_ filter The hc_ name filter was listed twice in the GLM5_NEXT protection block. Assisted-by: Claude Opus 5 * glm5-next: scope K-pool cache access to indexed operations * glm5-next: keep K-pool access in hybrid index memory * glm5-next: keep mHC graph builders model-local * glm5-next: mark only touched pools per ubatch --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
* add models backend check * t4-medium for faster build
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture test in the shared ggml-cuda sources evaluated to 0. Kernel bodies gated on the architecture therefore compiled to nothing, for example the q8_0 -> f16 dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to an empty body in host code. Report the newest architecture like the HIP backend does and exclude the NVIDIA-only features explicitly, as they are not usable on MUSA. Define it for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device compilation, which is also how nvcc behaves. Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and HIP defines it for every pass, so both forms select the same branch.
ggml-org#27773 adds the glm5-next arch without its rows in the Metal fusion baseline, so test-fusion --check fails on it. The rows come from test-fusion --record on an M5 Max, and --check passes 270/270.
…l-org#28381) * openvino: serve GET_ROWS on a weight view from the base Constant Resolve view_src when collecting weight Constants so a view over a quantized weight no longer becomes a dynamic typed Parameter, and fold the row offset of the view into the gather indices instead of slicing the dequantization subgraph. * openvino: lift the quantized GET_ROWS view rejection The supports_op rejection of a quantized src0 view with a nonzero offset keeps the vs0 GET_ROWS cases of ggml-org#28253 away from OpenVINO. The weight view now resolves to the base Constant with the row offset folded into the gather indices, so the rejection goes away.
…re plumbing (ggml-org#29582) * ui : type-safe API types, fetch helpers and download-ready models store plumbing Assisted-by: pi:GLM-5.3-Flash * ui : document the model list index pairing, fix an em-dash Assisted-by: pi:zai-org/GLM-5.3-Flash * Update tools/ui/src/lib/components/app/chat/index.ts Co-authored-by: Pascal <admin@serveurperso.com> --------- Co-authored-by: Pascal <admin@serveurperso.com>
…ml-org#27946) * ui : model id grammar for sidecars, quants and capability parsing Extend the shared model id parser with sidecar tokens (draft variants and auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add the tools capability to ModelCapabilities; the selector option row picks it up from the model's declared capabilities. Assisted-by: pi:GLM-5.3-Flash * ui : escape sidecar tokens in the regex alternation Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : Hugging Face Hub data layer Add HuggingFaceService and its constants/enums/types: GGUF repo search, file tree and model detail fetching, quant/sidecar filename analysis, shard-set collapsing and the llama.app catalog feed, plus an orgOf() helper on the model name utils. Assisted-by: pi:GLM-5.3-Flash * ui : strip provider tilde prefix from hub avatar urls Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : trim redundant comments in the HF data layer service Per review: drop JSDoc that restates the method name and inline comments that restate the code; keep only comments carrying non-obvious context. Assisted-by: pi:zai-org/GLM-5.3-Flash * ui : harden the HF data layer error typing, cover the helpers in tests Carries the HTTP status on retryable fetch errors instead of matching the message text. Marks expand-dependent catalog fields optional and documents the data/models index pairing. Adds table tests for the pure helpers. Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model memory-fit estimation Replace the raw runtime-memory estimate with the app's compatibility check: the smallest real Mac memory tier that fits a model file, budgeted as RAM x 0.75 minus fixed overhead with headroom on the file size. The constants move to lib; the unused runtime-memory estimate is dropped. browser-info's OS detection is exported for reuse. Assisted-by: pi:GLM-5.3-Flash * ui : cover the memory-fit and tool-use heuristics in tests Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model download pipeline Track HuggingFace downloads end to end: the server download/cancel endpoints, a status manager fed by the /models/sse download progress events, and a models-discover store holding the catalog and detail state for the discover view. Downloaded and in-flight entries are excluded from the loadable model list. Assisted-by: pi:GLM-5.3-Flash * ui : route sidecar tag lookup through the sidecars util, validate the paused list Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : shared model display primitives Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio order) out of ModelId and reuse it there, add the shared DialogConfirmDownload for destructive download actions, the discover org avatar with dark-mode inversion and the thin download progress bar, and rework ModelId badges to take thinking/tool-use support directly. Assisted-by: pi:GLM-5.3-Flash * ui : remember hub avatars that failed to load Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : render shared model row hints as native titles Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : fix badge guard for draft sidecars, keep parameter precision hasBadges now counts draft sidecar badges, so a sidecar-only model still renders. Billions keep one decimal for hub counts and stay bare for whole values. Avatar failures track the org instead of the instance, and the download progress bar no longer pulses while determinate. Assisted-by: pi:zai-org/GLM-5.3-Flash
) This commit adds an optional --add-bos token command line option to the run-org-model.py script. The motivation for this is that there are models, for example Gemma4, that explicitely set the add_bos value to true in llama-vocab.cpp even if the original model does not set this value to True. It would be nice to be able to force the models to agree on the bos token so that logit verification can proceed. Refs: ggml-org#21500
* cpu: accept BF16 in src1 of mul_mat ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in src1. The CPU backend refused that combination, so it was reported as unsupported on every backend and never compared against anything. Widen BF16 into the F32 work buffer, next to the existing packing of F32 into vec_dot_type. This is the arithmetic the Metal mat vec kernel already uses, both operands promoted to float and accumulated in float, so the two agree exactly rather than approximately. Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus three mul_mat cases with BF16 in src1. * vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16 supports_op only checked the src1 type for non contiguous tensors, so a contiguous BF16 src1 was accepted and the pipeline lookup asserted. The only BF16 src1 path is the BF16 x BF16 multiply, every other src0 type now reports the op as unsupported and the scheduler keeps it on the CPU. The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16 mat vec variants of the Metal backend, which land separately.
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
* llama: llama_prefetch_rows * llama: support row prefetch on Windows Apply the Windows port contributed by @praneshgo unchanged. Source: ggml-org#29599 (comment) * avoid exposing llama-mmap in model code, route via llama-impl * add windows check, only prefetch in lazy mode * cont : clean-up * cont : fix build * cont : clarify padding token for gemma4 --------- Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…l-org#29744) The fixture recycles its two blocks over 8 cache slots, so the fp16 error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU jobs of Models Backend. Two l-cycles keep every branch of the cycle loop and halve the error.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* support coerced array attributes * add tests
* CUDA: radix top-k for large row counts Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select, gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms. * CUDA: select the TOP_K implementation by shape Replace the nrows/ncols special case with the decision boundary from ggml-org#28547 (as implemented in ggml-org#29278): bitonic for short rows, radix select for several long rows, and DeviceTopK or CUB argsort for a single long row. The thresholds stay overridable at build time. Two refinements on top of that boundary: - bitonic stays in use for rows up to a padded 1024 while the rows fit in one wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of about a dozen launches that only amortizes over more rows - with DeviceTopK available, it handles up to two rows Radix select now processes rows in chunks so its scratch memory stays bounded, and the bitonic path keeps its chunking. HIP and MUSA keep their previous thresholds. Add perf cases around the bitonic/radix crossover to test-backend-ops. * CUDA: make top-k comments less verbose * CUDA: remove the TOP_K width limit from supports_op * CUDA: use DeviceTopK for single-row TOP_K if available * CUDA: avoid ncols overflow in the TOP_K bitonic check * CUDA: share the row chunking helper between argsort and top-k * CUDA: do the TOP_K radix blocks_per_row math in int64_t * CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD * CUDA: share one sort helper between the bitonic and CUB TOP_K paths * CUDA: update the TOP_K TODO, threshold and chunking comments * tests: add TOP_K cases that span several row chunks * CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment * CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX --------- Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com> Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Brings ggml-org/llama.cpp master (
de7fa0a) ontoclef-decision-port, which is tagrunanywhere-b10453.6.The Metal residency check stays gated on
MTLGPUFamilyApple6. That was the one local patch kept in the merge.tools/clef-demois removed. It still calledllama_batch_set_decision_order, and upstream moved that API tollama_batch_ext.Not in this PR
Text scores for d1-3B Q4 and d1-omni Q8 matched a local
llama-server/v1/systemonerun. Image scoring has not been run yet. GLiNER GGUF is not included. The SDK still calls the old decision API and is not updated here. Tagrunanywhere-b10453.7is the Windows ARM64 commit on top of.6; that cherry-pick is already upstream, so this branch is based on.6.Test plan
llama-clef-demogoneMade with Cursor
Summary by cubic
Syncs upstream ggml-org/llama.cpp master onto
clef-decision-port, therunanywhere-b10453.6tag.MTLGPUFamilyApple6; that is the only local patch kept in the merge.tools/clef-demois removed: it calledllama_batch_set_decision_order, and upstream moved that API tollama_batch_ext./v1/systemoneAPI andnimbledecision model support.llama-server/v1/systemonerun.runanywhere-b10453.7(Windows ARM64 on top of.6) is already upstream as a cherry-pick, so this branch is based on.6.Written for commit c6aa47c. Summary will update on new commits.