AMD 发布 Qwen3.8-27B DirectML ONNX 模型
AMD 在 Hugging Face 上发布了 Qwen3.8-27B 的 DirectML ONNX 版本,支持 FP16 视觉编码器和 INT4 文本解码器,适用于图像文本到文本任务。
证据来源
查看来源摘录
AMD amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx
查看来源摘录
amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx · Hugging Face Hugging Face (https://huggingface.co/) Models (https://huggingface.co/models) Datasets (https://huggingface.co/datasets) Spaces (https://huggingface.co/spaces) Buckets new (https://huggingface.co/storage) Docs (https://huggingface.co/docs) Enterprise (https://huggingface.co/enterprise) Pricing (https://huggingface.co/pricing) Website Tasks (https://huggingface.co/tasks) HuggingChat (https://huggingface.co/chat) Collections (https://huggingface.co/collections) Languages (https://huggingface.co/languages) Organizations (https://huggingface.co/organizations) Community Blog (https://huggingface.co/blog) Posts (https://huggingface.co/posts) Daily Papers (https://huggingface.co/papers) Hardware (https://huggingface.co/hardware) Learn (https://huggingface.co/learn) Discord (https://huggingface.co/join/discord) Forum (https://discuss.huggingface.co/) GitHub (https://github.com/huggingface) Solutions Team & Enterprise (https://huggingface.co/enterprise) Hugging Face PRO (https://huggingface.co/pro) Enterprise Support (https://huggingface.co/support) Inference Providers (https://huggingface.co/inference/models) Inference Endpoints (https://huggingface.co/inference-endpoints) Storage Buckets (https://huggingface.co/storage) Log In (https://huggingface.co/login) Sign Up (https://huggingface.co/join) (https://huggingface.co/amd) amd (https://huggingface.co/amd) / Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx) like 0 Follow AMD 3.15k Image-Text-to-Text (https://huggingface.co/models?pipeline_tag=image-text-to-text) ONNX (https://huggingface.co/models?library=onnx) English (https://huggingface.co/models?language=en) Chinese (https://huggingface.co/models?language=zh) onnxruntime-genai (https://huggingface.co/models?other=onnxruntime-genai) qwen3_5 (https://huggingface.co/models?other=qwen3_5) qwen (https://huggingface.co/models?other=qwen) qwen3.8 (https://huggingface.co/models?other=qwen3.8) directml (https://huggingface.co/models?other=directml) dml (https://huggingface.co/models?other=dml) int4 (https://huggingface.co/models?other=int4) vision (https://huggingface.co/models?other=vision) conversational (https://huggingface.co/models?other=conversational) License: apache-2.0 Model card (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx) Files Files and versions xet (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx/tree/main) Community 1 (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx/discussions) Copy to bucket new Qwen3.8-27B DirectML ONNX (FP16 vision/embedding, INT4 text) (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#qwen38-27b-directml-onnx-fp16-visionembedding-int4-text) Files (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#files) Usage (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#usage) Notes (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#notes) Export (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#export) (https://huggingface.co/amd/Qwen3.8-27B-fp16-ve-fp16-int4-k_quant-gs128-text-dml-onnx#qwen38-27b-directml-onnx-fp16-visionembedding-int4-text) Qwen3.8-27B DirectML ONNX (FP16 vision/embedding, INT4 text) ONNX Runtime GenAI package of Qwen/Qwen3.8-27B (https://huggingface.co/Qwen/Qwen3.8-27B) for DirectML. Subgraph Precision Notes Vision encoder (vision.onnx) FP16 Unquantized; LayerNormalization kept FP32 Token embedding (embedding.onnx) FP16 Unquantized Text decoder (text.onnx) INT4 Olive ModelBuilder k_quant, group/block size 128, accuracy_level=4. All 497 MatMulNBits weights are 4-bit. Activations
来自 星盘大模型百科