amd/GLM-5.3-Flash-Quark-MXFP4开源了新模型
根据官方来源,amd/GLM-5.3-Flash-Quark-MXFP4开源了新模型。详细信息请以原始来源为准。
证据来源
查看来源摘录
AMD amd/GLM-5.3-Flash-Quark-MXFP4 amd/GLM-5.3-Flash-Quark-MXFP4 https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4
查看来源摘录
amd/GLM-5.3-Flash-Quark-MXFP4 · Hugging Face Hugging Face (https://huggingface.co/) Models (https://huggingface.co/models) Datasets (https://huggingface.co/datasets) Spaces (https://huggingface.co/spaces) Buckets new (https://huggingface.co/storage) Docs (https://huggingface.co/docs) Enterprise (https://huggingface.co/enterprise) Pricing (https://huggingface.co/pricing) Website Tasks (https://huggingface.co/tasks) HuggingChat (https://huggingface.co/chat) Collections (https://huggingface.co/collections) Languages (https://huggingface.co/languages) Organizations (https://huggingface.co/organizations) Community Blog (https://huggingface.co/blog) Posts (https://huggingface.co/posts) Daily Papers (https://huggingface.co/papers) Hardware (https://huggingface.co/hardware) Learn (https://huggingface.co/learn) Discord (https://huggingface.co/join/discord) Forum (https://discuss.huggingface.co/) GitHub (https://github.com/huggingface) Solutions Team & Enterprise (https://huggingface.co/enterprise) Hugging Face PRO (https://huggingface.co/pro) Enterprise Support (https://huggingface.co/support) Inference Providers (https://huggingface.co/inference/models) Inference Endpoints (https://huggingface.co/inference-endpoints) Storage Buckets (https://huggingface.co/storage) Log In (https://huggingface.co/login) Sign Up (https://huggingface.co/join) (https://huggingface.co/amd) amd (https://huggingface.co/amd) / GLM-5.3-Flash-Quark-MXFP4 (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4) like 0 Follow AMD 3.2k Safetensors (https://huggingface.co/models?library=safetensors) glm5_next (https://huggingface.co/models?other=glm5_next) 8-bit precision (https://huggingface.co/models?other=8-bit) quark (https://huggingface.co/models?other=quark) License: mit Model card (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4) Files Files and versions xet (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4/tree/main) Community (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4/discussions) Copy to bucket new Model Overview (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#model-overview) Use with SGLang (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#use-with-sglang) Model Quantization (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#model-quantization) Use with SGLang (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#use-with-sglang) Deployment (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#deployment) Evaluation (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#evaluation) Accuracy (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#accuracy) Reproduction (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#reproduction) License (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#license) (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#model-overview) Model Overview Model Architecture: GLM-5.3-Flash Input: Text, Image, Video Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: PyTorch: Transformers: Operating System(s): Linux Inference Engine: Model Optimizer: AMD-Quark (https://quark.docs.amd.com/latest/index.html) (V0.12) Weight quantization: MOE-only (shared experts quantized), OCP MXFP4, Static Activation quantization: MOE-only, OCP MXFP4, Dynamic This model was built with GLM-5.3-Flash model by applying AMD-Quark (https://quark.docs.amd.com/latest/index.html) for MXFP4 quantization. (https://huggingface.co/amd/GLM-5.3-Flash-Quark-MXFP4#model-quantization) Model Quantization The model was quantized from zai-org/GLM-5.3-Flash (https://huggingface.co/zai-org/GLM-5.3-Flash) using AMD-Quark (https://quark.docs.amd.com/latest/index.html) . The weights and activations are quantized to MXFP4. Quantization scripts: from quark.torch import LLMTemplate, ModelQuantizer EXCLUDE = [ "*self_attn*", "*mlp.gate", "*mlp.gate_proj", "*mlp.up_proj", "*mlp.down_proj", # dense MLP (layers 0-2) "*visual*", "*lm_head*", "*embed*", "model.l
来自 星盘大模型百科