返回模型
说明文档
ChatGLM-6B + ONNX
该模型是从 ChatGLM-6b 导出的,采用 int8 量化并针对 ONNXRuntime 推理进行了优化。导出代码在这个仓库中。
基于 ONNXRuntime 的推理代码已随模型上传。安装依赖并运行 streamlit run web-ui.py 即可开始对话。目前 MatMulInteger(用于 u8s8 数据类型)和 DynamicQuantizeLinear 算子仅支持 CPU。支持 Neon 的 Arm64(Apple M1/M2)应该有不错的速度。
安装依赖并运行 streamlit run web-ui.py 预览模型效果。由于 ONNXRuntime 算子支持问题,目前仅能够使用 CPU 进行推理,在 Arm64 (Apple M1/M2) 上有可观的速度。具体的 ONNX 导出代码在这个仓库中。
使用方法
使用 git-lfs 克隆:
git lfs clone https://huggingface.co/K024/ChatGLM-6b-onnx-u8s8
cd ChatGLM-6b-onnx-u8s8
pip install -r requirements.txt
streamlit run web-ui.py
或者使用 huggingface_hub Python 客户端库下载仓库快照:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="K024/ChatGLM-6b-onnx-u8s8", local_dir="./ChatGLM-6b-onnx-u8s8")
代码基于 MIT 许可证发布。
模型权重基于与 ChatGLM-6b 相同的许可证发布,详见 MODEL LICENSE。
K024/ChatGLM-6b-onnx-u8s8
作者 K024
↓ 0
♥ 11
创建时间: 2023-04-28 07:52:54+00:00
更新时间: 2023-05-16 07:48:13+00:00
在 Hugging Face 上查看文件 (16)
.gitattributes
.gitignore
README.md
chatglm-6b-int8-onnx-merged/chatglm-6b-int8.onnx
ONNX
chatglm-6b-int8-onnx-merged/model_weights_0.bin
chatglm-6b-int8-onnx-merged/model_weights_1.bin
chatglm-6b-int8-onnx-merged/model_weights_2.bin
chatglm-6b-int8-onnx-merged/model_weights_3.bin
chatglm-6b-int8-onnx-merged/model_weights_4.bin
chatglm-6b-int8-onnx-merged/model_weights_5.bin
chatglm-6b-int8-onnx-merged/model_weights_6.bin
chatglm-6b-int8-onnx-merged/sentencepiece.model
model.py
requirements.txt
tokenizer.py
web-ui.py