返回模型
说明文档
Qwen3-4B-ONNX-INT4-CPU
注意: 这是一个非官方版本,仅用于测试和开发目的。
模型转换
本指南演示如何使用 Microsoft Olive 和 Microsoft ONNXRuntime GenAI 将 Qwen3-4B 模型转换为 ONNX 格式。
前置条件
请确保已安装以下工具:
- CMake 3.31+
- Transformers 4.51+
- Microsoft Olive(开发版本)
- Microsoft ONNXRuntime GenAI(开发版本)
安装步骤
# 更新 transformers 库
pip install transformers -U
# 安装 Microsoft Olive
pip install git+https://github.com/microsoft/Olive.git
# 安装 Microsoft ONNXRuntime GenAI
git clone https://github.com/microsoft/onnxruntime-genai
cd onnxruntime-genai && python build.py --config Release
注意: 如果您尚未安装 CMake,请先安装后再继续。
转换命令
# 使用 Microsoft Olive 转换模型
olive auto-opt \
--model_name_or_path {Qwen3-4B_PATH} \
--device cpu \
--provider CPUExecutionProvider \
--use_model_builder \
--precision int4 \
--output_path {Your_Qwen3-4B_ONNX_Output_Path} \
--log_level 1
推理
Qwen3 支持两种推理模式,具有不同的参数配置:
思考模式
当您希望模型展示其推理过程时:
- 参数: Temperature=0.6, TopP=0.95, TopK=20, MinP=0.0
- 对话模板:
<|im_start|>user\n/think {input}<|im_end|><|im_start|>assistant\n
非思考模式
用于直接响应,不展示推理过程:
- 参数: Temperature=0.7, TopP=0.8, TopK=20, MinP=0.0
- 对话模板:
<|im_start|>user\n/no_think {input}<|im_end|><|im_start|>assistant\n
Python 示例
import onnxruntime_genai as og
import json
# 设置模型路径
model_folder = "Your_Qwen3-4B_ONNX_Path"
# 初始化模型和分词器
model = og.Model(model_folder)
tokenizer = og.Tokenizer(model)
tokenizer_stream = tokenizer.create_stream()
# 思考模式配置
search_options = {
'temperature': 0.6,
'top_p': 0.95,
'top_k': 20,
'max_length': 32768,
'repetition_penalty': 1
}
chat_template = "<|im_start|>user\n/think {input}<|im_end|><|im_start|>assistant\n"
text = 'What is the derivative of x^2?'
# 非思考模式的替代配置
# search_options = {
# 'temperature': 0.7,
# 'top_p': 0.8,
# 'top_k': 20,
# 'max_length': 4096,
# 'repetition_penalty': 1
# }
# chat_template = "<|im_start|>user\n/no_think {input}<|im_end|><|im_start|>assistant\n"
# text = 'Can you introduce yourself?'
# 准备提示词并生成响应
prompt = chat_template.format(input=text)
input_tokens = tokenizer.encode(prompt)
params = og.GeneratorParams(model)
params.set_search_options(**search_options)
generator = og.Generator(model, params)
generator.append_tokens(input_tokens)
# 生成并流式输出响应
while not generator.is_done():
generator.generate_next_token()
new_token = generator.get_next_tokens()[0]
print(tokenizer_stream.decode(new_token), end='', flush=True)
lokinfey/Qwen3-4B-ONNX-INT4-CPU
作者 lokinfey
↓ 1
♥ 0
创建时间: 2025-06-29 04:11:05+00:00
更新时间: 2025-06-29 05:26:10+00:00
在 Hugging Face 上查看文件 (14)
.gitattributes
README.md
added_tokens.json
chat_template.jinja
config.json
genai_config.json
generation_config.json
merges.txt
model.onnx
ONNX
model.onnx.data
special_tokens_map.json
tokenizer.json
tokenizer_config.json
vocab.json