Skip to content

port: clef decision-model support + llama-clef-demo (ABI-compatible with b10453.4) - #1

Closed
Siddhesh2377 wants to merge 12 commits into
RunanywhereAI:runanywhere-b10453-maple-prismfrom
Siddhesh2377:clef-decision-port
Closed

Siddhesh2377 wants to merge 12 commits into
RunanywhereAI:runanywhere-b10453-maple-prismfrom
Siddhesh2377:clef-decision-port

Conversation

@Siddhesh2377

@Siddhesh2377 Siddhesh2377 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

What

Ports the clef decision-model architecture onto runanywhere-b10453.4 and wires the joint decision head end to end, plus a CLI that exercises it.

A clef GGUF (general.architecture = clef) attaches a joint decision head: one forward pass over a prompt carrying a state + every typed question + every option scores all options at once. The head reads per-token span tags to pool the right runs, and the per-option scores come back as rows of the embeddings output.

Commits

  1. clef model-side port — LLM_ARCH_CLEF registration, arch/KV/tensor tables, llama_model_clef (Qwen3.5 backbone + joint head), rope/memory/graph routing, model saver opt-out, arch test, conversion/clef.py + gguf-py metadata.
  2. compat fixes for this base — arch-table duplicate removal, span array inits, decision-order enum in the public header, the missing build_qkv overload the head calls.
  3. llama_batch_set_decision_order — span array stays NULL after llama_batch_init and is allocated on first setter call. NULL means "no spans", which is what keeps plain generation on a decision model on the 0.0f status path instead of the all-zeros NaN path.
  4. llama-clef-demo — a CLI that runs a real joint decision: spans tagged, one llama_encode, one score per option, softmax → choice/confidence. Named demos + generic --question/--option.

Verification

Built with -DLLAMA_BUILD_TESTS=ON and run against Clef-Flash-Q4_K_M.gguf:

  • test-llama-archs lists clef; model loads all 587 tensors with no arch errors.
  • llama-clef-demo --demo → A (sqlite-vec) at p=0.92, exit 0.
  • Plain generation (no spans) still takes the clean path.

Notes

  • Model-side fidelity: src/models/clef.cpp and conversion/clef.py are the same content as the upstream clef support, adapted to this base's batch model.
  • The server-side decision endpoints are intentionally not ported; this base has no extended-batch framework and the SDK consumes the head through its own engine op.
  • Maple entries untouched; clef is a dense Qwen3.5-9B, no interaction expected.

Summary by cubic

Ports the clef decision-model architecture onto runanywhere-b10453.4 so a clef GGUF scores all options in one forward pass. Plain generation on a decision model keeps its previous behavior: batches without tagged spans take the clean path and the head scores nothing.

What's included

  • Registers LLM_ARCH_CLEF with a Qwen3.5 backbone plus joint decision head, model mapping, rope/memory/graph routing, arch test, conversion/clef.py, and gguf-py metadata.
  • Adds llama_batch_set_decision_order to tag token spans; the span array stays NULL (meaning "no spans") until the first setter call, so batches without spans keep the clean 0.0f status path instead of the invalid-input NaN path.
  • Adds llama-clef-demo, a CLI that tags question/option spans, runs one llama_encode, and reads one score per option back from the embeddings output with softmax → choice/confidence, including named demos and generic --question/--option input.

Not included

  • Server-side decision endpoints, which need the extended-batch framework this base lacks.
  • Model saver support for clef — the head tensors are not saved. Maple entries are untouched.

Written for commit adcfbd7. Summary will update on new commits.

Review in cubic

ngxson and others added 12 commits October 3, 2026 12:34
Scoped port of the clef decision-model architecture onto runanywhere-b10453.4. Registers LLM_ARCH_CLEF ('clef'):
arch/KV/tensor tables, llama_model_clef (Qwen3.5 backbone + joint
decision head), model mapping, rope/memory/graph-node routing, model
saver opt-out, arch-test entry, conversion/clef.py + gguf-py metadata.

Deliberately excluded (needs the newer batch/common/server
framework refactor absent from this fork): extended-batch
decision_order plumbing, common decision-type helpers, server decision
endpoints, extended batch span setter. QWEN4EXP drift from
the same newer window also excluded; Maple entries untouched.
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…init, ext enum, graph build_qkv overload)

Backports the compile-time pieces that the b10453 base lacks. common/server-side decision parts intentionally
skipped: they depend on the extended batch framework which does not
exist in this base.
…oint head)

Old base has no extended batch framework; carry spans on the classic
llama_batch struct (calloc'd by llama_batch_init, freed by
llama_batch_free) and copy them in ubatch_add. NULL (default) =
no spans, same as before.
Runs a joint-head decision end to end: question + option spans tagged via
llama_decision_order, one llama_encode pass, one score per option read from
the embeddings output, softmax -> choice + confidence.

Includes named demo cases (db, privacy, privacy-reversed, privacy-subtle)
and generic --question/--option input with choice|score|noul types.
Descriptive wording only; no functional change.
…unset

The span array stays NULL after llama_batch_init and is allocated on the
first setter call. This restores the NULL-means-absent contract: batches
without spans take the 0.0f status path instead of the all-zeros NaN path,
so plain generation on a decision model is not poisoned. The decision
order enum moves to the public header next to the batch field it types.
@Siddhesh2377

Copy link
Copy Markdown
Collaborator Author

Retargeting: the branch now lives in this repo (RunanywhereAI/llama.cpp:clef-decision-port). Opening the in-repo PR next.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants