返回模型
说明文档
ImageGPT(小型模型)
ImageGPT (iGPT) 模型在 ImageNet ILSVRC 2012(1400万张图像,21,843个类别)上以 32x32 分辨率预训练。该模型由 Chen 等人在论文 Generative Pretraining from Pixels 中提出,并首次在 此仓库 中发布。另请参阅官方 博客文章。
模型描述
ImageGPT (iGPT) 是一个 Transformer 解码器模型(类 GPT),以自监督方式在大量图像集合(即 ImageNet-21k)上预训练,分辨率为 32x32 像素。
该模型的目标很简单:根据前面的像素值预测下一个像素值。
通过预训练,模型学习到图像的内部表示,可用于:
- 提取对下游任务有用的特征:可以使用 ImageGPT 生成固定的图像特征,以训练线性模型(如 sklearn 逻辑回归模型或 SVM)。这也被称为"线性探测"(linear probing)。
- 执行(非)条件图像生成。
预期用途与限制
您可以将原始模型用作特征提取器或(非)条件图像生成。
如何使用
以下是如何将此模型用作特征提取器:
from transformers import AutoFeatureExtractor
from onnxruntime import InferenceSession
from datasets import load_dataset
# 加载图像
dataset = load_dataset("huggingface/cats-image")
image = dataset["test"]["image"][0]
# 加载模型
feature_extractor = AutoFeatureExtractor.from_pretrained("openai/imagegpt-small")
session = InferenceSession("model/model.onnx")
# ONNX Runtime 期望 NumPy 数组作为输入
inputs = feature_extractor(image, return_tensors="np")
outputs = session.run(output_names=["last_hidden_state"], input_feed=dict(inputs))
或者,您可以使用带有分类头的模型,该模型返回 logits:
from transformers import AutoFeatureExtractor
from onnxruntime import InferenceSession
from datasets import load_dataset
# 加载图像
dataset = load_dataset("huggingface/cats-image")
image = dataset["test"]["image"][0]
# 加载模型
feature_extractor = AutoFeatureExtractor.from_pretrained("openai/imagegpt-small")
session = InferenceSession("model/model_classification.onnx")
# ONNX Runtime 期望 NumPy 数组作为输入
inputs = feature_extractor(image, return_tensors="np")
outputs = session.run(output_names=["logits"], input_feed=dict(inputs))
原始实现
访问 此链接 查看原始实现。
训练数据
ImageGPT 模型在 ImageNet-21k 上预训练,该数据集包含 1400 万张图像和 21k 个类别。
训练过程
预处理
图像首先被调整/缩放到相同分辨率(32x32),并在 RGB 通道上进行归一化。接下来,进行颜色聚类。这意味着每个像素被转换为 512 个可能的聚类值之一。这样,最终得到 32x32 = 1024 个像素值序列,而不是 32x32x3 = 3072,后者对于基于 Transformer 的模型来说过于庞大。
预训练
训练详情可在论文 v2 版本的第 3.4 节中找到。
评估结果
关于多个图像分类基准的评估结果,请参阅原始论文。
BibTeX 条目和引用信息
@InProceedings{pmlr-v119-chen20s,
title = {Generative Pretraining From Pixels},
author = {Chen, Mark and Radford, Alec and Child, Rewon and Wu, Jeffrey and Jun, Heewoo and Luan, David and Sutskever, Ilya},
booktitle = {Proceedings of the 37th International Conference on Machine Learning},
pages = {1691--1703},
year = {2020},
editor = {III, Hal Daumé and Singh, Aarti},
volume = {119},
series = {Proceedings of Machine Learning Research},
month = {13--18 Jul},
publisher = {PMLR},
pdf = {http://proceedings.mlr.press/v119/chen20s/chen20s.pdf},
url = {https://proceedings.mlr.press/v119/chen20s.html
}
@inproceedings{deng2009imagenet,
title={Imagenet: A large-scale hierarchical image database},
author={Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Li, Kai and Fei-Fei, Li},
booktitle={2009 IEEE conference on computer vision and pattern recognition},
pages={248--255},
year={2009},
organization={Ieee}
}
OWG/imagegpt-small
作者 OWG
↓ 0
♥ 0
创建时间: 2022-10-27 11:52:39+00:00
更新时间: 2022-10-27 13:10:17+00:00
在 Hugging Face 上查看文件 (4)
.gitattributes
README.md
model/model.onnx
ONNX
model/model_classification.onnx
ONNX