amd/GLM-5.3-Quark-MXFP4-AttnFP8开源了新模型
根据官方来源,amd/GLM-5.3-Quark-MXFP4-AttnFP8开源了新模型。详细信息请以原始来源为准。
证据来源
查看来源摘录
AMD amd/GLM-5.3-Quark-MXFP4-AttnFP8 amd/GLM-5.3-Quark-MXFP4-AttnFP8 https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8
查看来源摘录
amd/GLM-5.3-Quark-MXFP4-AttnFP8 · Hugging Face Hugging Face (https://huggingface.co/) Models (https://huggingface.co/models) Datasets (https://huggingface.co/datasets) Spaces (https://huggingface.co/spaces) Buckets new (https://huggingface.co/storage) Docs (https://huggingface.co/docs) Enterprise (https://huggingface.co/enterprise) Pricing (https://huggingface.co/pricing) Website Tasks (https://huggingface.co/tasks) HuggingChat (https://huggingface.co/chat) Collections (https://huggingface.co/collections) Languages (https://huggingface.co/languages) Organizations (https://huggingface.co/organizations) Community Blog (https://huggingface.co/blog) Posts (https://huggingface.co/posts) Daily Papers (https://huggingface.co/papers) Hardware (https://huggingface.co/hardware) Learn (https://huggingface.co/learn) Discord (https://huggingface.co/join/discord) Forum (https://discuss.huggingface.co/) GitHub (https://github.com/huggingface) Solutions Team & Enterprise (https://huggingface.co/enterprise) Hugging Face PRO (https://huggingface.co/pro) Enterprise Support (https://huggingface.co/support) Inference Providers (https://huggingface.co/inference/models) Inference Endpoints (https://huggingface.co/inference-endpoints) Storage Buckets (https://huggingface.co/storage) Log In (https://huggingface.co/login) Sign Up (https://huggingface.co/join) (https://huggingface.co/amd) amd (https://huggingface.co/amd) / GLM-5.3-Quark-MXFP4-AttnFP8 (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8) like 0 Follow AMD 3.2k Safetensors (https://huggingface.co/models?library=safetensors) glm_moe_dsa (https://huggingface.co/models?other=glm_moe_dsa) Eval Results (https://huggingface.co/models?other=eval-results) 8-bit precision (https://huggingface.co/models?other=8-bit) quark (https://huggingface.co/models?other=quark) License: mit Model card (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8) Files Files and versions xet (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8/tree/main) Community 1 (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8/discussions) Copy to bucket new Model Overview (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#model-overview) Use with vLLM (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#use-with-vllm) Model Quantization (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#model-quantization) Use with vLLM (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#use-with-vllm) Deployment (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#deployment) Evaluation (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#evaluation) Accuracy (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#accuracy) Reproduction (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#reproduction) License (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#license) (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#model-overview) Model Overview Model Architecture: GlmMoeDsaForCausalLM Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.1.1 PyTorch: 2.10.0+rocm7.1.1.lw.gitd9556b05 Transformers: 5.15.1 Operating System(s): Linux Inference Engine: vLLM (https://docs.vllm.ai/en/latest/) Model Optimizer: AMD-Quark (https://quark.docs.amd.com/latest/index.html) (V0.12) Quantized layers: self_attn, router experts, shared_experts self_attn: FP8 per-block router experts and shared_experts: MXFP4 This model was built with GLM-5.3 model by applying AMD-Quark (https://quark.docs.amd.com/latest/index.html) for MXFP4 quantization. (https://huggingface.co/amd/GLM-5.3-Quark-MXFP4-AttnFP8#model-quantization) Model Quantization The model was quantized from zai-org/GLM-5.3 (https://huggingface.co/zai-org/GLM-5.3) using AMD-Quark (https://quark.docs.amd.com/latest/index.html) . The weights and activations are quantized to MXFP4. Quantization scripts: cd Quark/examples/torch/language_modeling/llm_ptq/ python quantize
来自 星盘大模型百科