https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8
https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8
I was told it will require manual quantization due to size, but I would highly appreciate it if you could do so!
We will do it as soon https://github.com/ggml-org/llama.cpp/pull/27773 is merged
This is the full GLM-5.3, not flash. I believe it works just the same as 5.2 in llama.cpp
Oh interesting your requested model is of archidecture GlmMoeDsaForCausalLM and not Glm5NextForConditionalGeneration so we can give it a try.
Oh interesting your requested model is of archidecture
GlmMoeDsaForCausalLMand notGlm5NextForConditionalGenerationso we can give it a try.
Awesome, do you have a rough estimate when they will be finished?
For anyone else who might have been waiting on this like me, someone else uploaded GGUF quants, Q4_K_M & IQ2_M: https://huggingface.co/softwareweaver/GLM-5.3-Uncensored-GGUF