ONNX 模型库
返回模型

说明文档

ImageGPT(小型模型)

ImageGPT (iGPT) 模型在 ImageNet ILSVRC 2012(1400万张图像,21,843个类别)上以 32x32 分辨率预训练。该模型由 Chen 等人在论文 Generative Pretraining from Pixels 中提出,并首次在 此仓库 中发布。另请参阅官方 博客文章。

模型描述

ImageGPT (iGPT) 是一个 Transformer 解码器模型(类 GPT),以自监督方式在大量图像集合(即 ImageNet-21k)上预训练,分辨率为 32x32 像素。

该模型的目标很简单:根据前面的像素值预测下一个像素值。

通过预训练,模型学习到图像的内部表示,可用于:

  • 提取对下游任务有用的特征:可以使用 ImageGPT 生成固定的图像特征,以训练线性模型(如 sklearn 逻辑回归模型或 SVM)。这也被称为"线性探测"(linear probing)。
  • 执行(非)条件图像生成。

预期用途与限制

您可以将原始模型用作特征提取器或(非)条件图像生成。

如何使用

以下是如何将此模型用作特征提取器:

from transformers import AutoFeatureExtractor
from onnxruntime import InferenceSession
from datasets import load_dataset

# 加载图像
dataset = load_dataset("huggingface/cats-image")
image = dataset["test"]["image"][0]

# 加载模型
feature_extractor = AutoFeatureExtractor.from_pretrained("openai/imagegpt-small")
session = InferenceSession("model/model.onnx")

# ONNX Runtime 期望 NumPy 数组作为输入
inputs = feature_extractor(image, return_tensors="np")
outputs = session.run(output_names=["last_hidden_state"], input_feed=dict(inputs))

或者,您可以使用带有分类头的模型,该模型返回 logits:

from transformers import AutoFeatureExtractor
from onnxruntime import InferenceSession
from datasets import load_dataset

# 加载图像
dataset = load_dataset("huggingface/cats-image")
image = dataset["test"]["image"][0]

# 加载模型
feature_extractor = AutoFeatureExtractor.from_pretrained("openai/imagegpt-small")
session = InferenceSession("model/model_classification.onnx")

# ONNX Runtime 期望 NumPy 数组作为输入
inputs = feature_extractor(image, return_tensors="np")
outputs = session.run(output_names=["logits"], input_feed=dict(inputs))

原始实现

访问 此链接 查看原始实现。

训练数据

ImageGPT 模型在 ImageNet-21k 上预训练,该数据集包含 1400 万张图像和 21k 个类别。

训练过程

预处理

图像首先被调整/缩放到相同分辨率(32x32),并在 RGB 通道上进行归一化。接下来,进行颜色聚类。这意味着每个像素被转换为 512 个可能的聚类值之一。这样,最终得到 32x32 = 1024 个像素值序列,而不是 32x32x3 = 3072,后者对于基于 Transformer 的模型来说过于庞大。

预训练

训练详情可在论文 v2 版本的第 3.4 节中找到。

评估结果

关于多个图像分类基准的评估结果,请参阅原始论文。

BibTeX 条目和引用信息

@InProceedings{pmlr-v119-chen20s,
  title = 	 {Generative Pretraining From Pixels},
  author =       {Chen, Mark and Radford, Alec and Child, Rewon and Wu, Jeffrey and Jun, Heewoo and Luan, David and Sutskever, Ilya},
  booktitle = 	 {Proceedings of the 37th International Conference on Machine Learning},
  pages = 	 {1691--1703},
  year = 	 {2020},
  editor = 	 {III, Hal Daumé and Singh, Aarti},
  volume = 	 {119},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--18 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v119/chen20s/chen20s.pdf},
  url = 	 {https://proceedings.mlr.press/v119/chen20s.html
}
@inproceedings{deng2009imagenet,
  title={Imagenet: A large-scale hierarchical image database},
  author={Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Li, Kai and Fei-Fei, Li},
  booktitle={2009 IEEE conference on computer vision and pattern recognition},
  pages={248--255},
  year={2009},
  organization={Ieee}
}

OWG/imagegpt-small

作者 OWG

↓ 0 ♥ 0

创建时间: 2022-10-27 11:52:39+00:00

更新时间: 2022-10-27 13:10:17+00:00

在 Hugging Face 上查看

文件 (4)

.gitattributes
README.md
model/model.onnx ONNX
model/model_classification.onnx ONNX