-
Notifications
You must be signed in to change notification settings - Fork 1.1k
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[Question]将qwen3.5-27B转换为torch dist后,再转换回hf格式,使用sglang部署,在tp=4/8的情况下模型会输出且只输出大量‘!’
questionFurther information is requestedFurther information is requestedStatus: Open.#2222 In THUDM/slime;[Bug] GLM5.2 模型转换的时候报错
bugSomething isn't workingSomething isn't workingStatus: Open.#2215 In THUDM/slime;- Status: Open.#2214 In THUDM/slime;
[Question] Do we support the SAO method used in the training of GLM-5.2
questionFurther information is requestedFurther information is requestedStatus: Open.#2212 In THUDM/slime;- Status: Open.#2209 In THUDM/slime;
[Question] Do we support general on-policy distillation?
questionFurther information is requestedFurther information is requestedStatus: Open.#2202 In THUDM/slime;- Status: Open.#2201 In THUDM/slime;
[Question] qwen3.6-35B-A3b训练的时候只保存了文本权重,多模态权重没有保存,怎么办
questionFurther information is requestedFurther information is requestedStatus: Open.#2194 In THUDM/slime;[Bug] Colocate weight update fails with torch_memory_saver/offload because PyTorch CUDA IPC _share_cuda_ raises cudaErrorInvalidValue
bugSomething isn't workingSomething isn't workingStatus: Open.#2188 In THUDM/slime;[Bug]
offload_traintrain worker dies on CUDA 13 withlibcudart.so.12: cannot open shared object file— the torch_memory_saverLD_PRELOAD.sois chosen by filename, not by CUDA runtimebugSomething isn't workingSomething isn't workingStatus: Open.#2186 In THUDM/slime;[Bug] 过采样关停的时候judge不会被杀掉
bugSomething isn't workingSomething isn't workingStatus: Open.#2176 In THUDM/slime;[Bug] Using q instead of normalized q in Megatron's DSA MLA indexer for GLM 5 models
bugSomething isn't workingSomething isn't workingStatus: Open.#2165 In THUDM/slime;