说明文档
模型概述
这是一个经过微调的 xlm-roberta 模型,用于恢复标点符号、正确大小写(首字母大写)以及检测47种语言的句子边界(句号)。
使用方法
如果您只是想体验一下这个模型,本页面的小部件就足够了。如果您想离线使用该模型,以下代码片段展示了如何通过封装库(我编写的,可从 PyPI 获取)以及手动使用(使用本仓库中的 ONNX 和 SentencePiece 模型)两种方式来使用该模型。
通过 punctuators 包使用
<details>
<summary>点击查看使用封装库的方法</summary>
使用该模型最简单的方式是安装 punctuators:
$ pip install punctuators
但这只是一个 ONNX 和 SentencePiece 模型,因此您可以按自己的方式运行它。
punctuators API 的输入是一个字符串列表(批次)。每个字符串将被添加标点、正确大小写,并在预测的句号处分割。因此,输出将是一个字符串列表的列表:每个输入文本对应一个分割后的句子列表。要禁用句号分割功能,请使用 m.infer(texts, apply_sbd=False)。此时输出将是一个字符串列表:每个输入文本对应一个添加了标点且正确大小写的字符串。
<details open>
<summary>使用示例</summary>
from typing import List
from punctuators.models import PunctCapSegModelONNX
m: PunctCapSegModelONNX = PunctCapSegModelONNX.from_pretrained(
\"1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase\"
)
input_texts: List[str] = [
\"hola mundo cómo estás estamos bajo el sol y hace mucho calor santa coloma abre los huertos urbanos a las escuelas de la ciudad\",
\"hello friend how's it going it's snowing outside right now in connecticut a large storm is moving in\",
\"未來疫苗將有望覆蓋3歲以上全年齡段美國與北約軍隊已全部撤離還有鐵路公路在內的各項基建的來源都將枯竭\",
\"በባለፈው ሳምንት ኢትዮጵያ ከሶማሊያ 3 ሺህ ወታደሮቿንም እንዳስወጣች የሶማሊያው ዳልሳን ሬድዮ ዘግቦ ነበር ጸጥታ ሃይሉና ህዝቡ ተቀናጅቶ በመስራቱ በመዲናዋ ላይ የታቀደው የጥፋት ሴራ ከሽፏል\",
\"こんにちは友人\" \"調子はどう\" \"今日は雨の日でしたね\" \"乾いた状態を保つために一日中室内で過ごしました\",
\"hallo freund wie geht's es war heute ein regnerischer tag nicht wahr ich verbrachte den tag drinnen um trocken zu bleiben\",
\"हैलो दोस्त ये कैसा चल रहा है आज बारिश का दिन था न मैंने सूखा रहने के लिए दिन घर के अंदर बिताया\",
\"كيف تجري الامور كان يومًا ممطرًا اليوم أليس كذلك قضيت اليوم في الداخل لأظل جافًا\",
]
results: List[List[str]] = m.infer(
texts=input_texts, apply_sbd=True,
)
for input_text, output_texts in zip(input_texts, results):
print(f\"Input: {input_text}\")
print(f\"Outputs:\")
for text in output_texts:
print(f\"\t{text}\")
print()
</details>
<details open>
<summary>预期输出</summary>
Input: hola mundo cómo estás estamos bajo el sol y hace mucho calor santa coloma abre los huertos urbanos a las escuelas de la ciudad
Outputs:
Hola mundo, ¿cómo estás?
Estamos bajo el sol y hace mucho calor.
Santa Coloma abre los huertos urbanos a las escuelas de la ciudad.
Input: hello friend how's it going it's snowing outside right now in connecticut a large storm is moving in
Outputs:
Hello friend, how's it going?
It's snowing outside right now.
In Connecticut, a large storm is moving in.
Input: 未來疫苗將有望覆蓋3歲以上全年齡段美國與北約軍隊已全部撤離還有鐵路公路在內的各項基建的來源都將枯竭
Outputs:
未來,疫苗將有望覆蓋3歲以上全年齡段。
美國與北約軍隊已全部撤離。
還有,鐵路,公路在內的各項基建的來源都將枯竭。
Input: በባለፈው ሳምንት ኢትዮጵያ ከሶማሊያ 3 ሺህ ወታደሮቿንም እንዳስወጣች የሶማሊያው ዳልሳን ሬድዮ ዘግቦ ነበር ጸጥታ ሃይሉና ህዝቡ ተቀናጅቶ በመስራቱ በመዲናዋ ላይ የታቀደው የጥፋት ሴራ ከሽፏል
Outputs:
በባለፈው ሳምንት ኢትዮጵያ ከሶማሊያ 3 ሺህ ወታደሮቿንም እንዳስወጣች የሶማሊያው ዳልሳን ሬድዮ ዘግቦ ነበር።
ጸጥታ ሃይሉና ህዝቡ ተቀናጅቶ በመስራቱ በመዲናዋ ላይ የታቀደው የጥፋት ሴራ ከሽፏል።
Input: こんにちは友人調子はどう今日は雨の日でしたね乾いた状態を保つために一日中室内で過ごしました
Outputs:
こんにちは、友人、調子はどう?
今日は雨の日でしたね。
乾いた状態を保つために、一日中、室内で過ごしました。
Input: hallo freund wie geht's es war heute ein regnerischer tag nicht wahr ich verbrachte den tag drinnen um trocken zu bleiben
Outputs:
Hallo Freund, wie geht's?
Es war heute ein regnerischer Tag, nicht wahr?
Ich verbrachte den Tag drinnen, um trocken zu bleiben.
Input: हैलो दोस्त ये कैसा चल रहा है आज बारिश का दिन था न मैंने सूखा रहने के लिए दिन घर के अंदर बिताया
Outputs:
हैलो दोस्त, ये कैसा चल रहा है?
आज बारिश का दिन था न, मैंने सूखा रहने के लिए दिन घर के अंदर बिताया।
Input: كيف تجري الامور كان يومًا ممطرًا اليوم أليس كذلك قضيت اليوم في الداخل لأظل جافًا
Outputs:
كيف تجري الامور؟
كان يومًا ممطرًا اليوم، أليس كذلك؟
قضيت اليوم في الداخل لأظل جافًا.
</details>
</details>
手动使用
如果您想在没有封装库的情况下使用 ONNX 和 SP 模型,请参阅以下示例。
<details>
<summary>点击查看手动使用方法</summary>
from typing import List
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from omegaconf import OmegaConf
from sentencepiece import SentencePieceProcessor
# 从 HF hub 下载模型。注意:如需清理,您可以在 HF 缓存目录中找到这些文件
spe_path = hf_hub_download(repo_id=\"1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase\", filename=\"sp.model\")
onnx_path = hf_hub_download(repo_id=\"1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase\", filename=\"model.onnx\")
config_path = hf_hub_download(
repo_id=\"1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase\", filename=\"config.yaml\"
)
# 加载 SP 模型
tokenizer: SentencePieceProcessor = SentencePieceProcessor(spe_path) # noqa
# 加载 ONNX 图
ort_session: ort.InferenceSession = ort.InferenceSession(onnx_path)
# 加载包含标签等的模型配置
config = OmegaConf.load(config_path)
# 每个 subtoken 前可能的分类标签
pre_labels: List[str] = config.pre_labels
# 每个 subtoken 后可能的分类标签
post_labels: List[str] = config.post_labels
# 表示"不预测任何内容"的特殊类别
null_token = config.get(\"null_token\", \"<NULL>\")
# 表示"此 subtoken 中的所有字符都以句点结尾"的特殊类别,例如 "am" -> "a.m."
acronym_token = config.get(\"acronym_token\", \"<ACRONYM>\")
# 本示例中未使用,但如果您的序列超过此值,则需要将其折叠到多个输入中
max_len = config.max_length
# 仅供参考,该图没有特定语言的行为
languages: List[str] = config.languages
# 编码一些输入文本,添加 BOS + EOS
input_text = \"hola mundo cómo estás estamos bajo el sol y hace mucho calor santa coloma abre los huertos urbanos a las escuelas de la ciudad\"
input_ids = [tokenizer.bos_id()] + tokenizer.EncodeAsIds(input_text) + [tokenizer.eos_id()]
# 创建一个形状为 [B, T] 的 numpy 数组,作为图的输入。
# 注意,我们不向图传递长度;如果您使用批次,填充应为 tokenizer.pad_id(),并且
# 图的注意力机制将忽略 pad_id(),而无需显式的序列长度。
input_ids_arr: np.array = np.array([input_ids])
# 运行图,获取所有分析的输出
pre_preds, post_preds, cap_preds, sbd_preds = ort_session.run(None, {\"input_ids\": input_ids_arr})
# 去除批次维度并转换为列表
pre_preds = pre_preds[0].tolist()
post_preds = post_preds[0].tolist()
cap_preds = cap_preds[0].tolist()
sbd_preds = sbd_preds[0].tolist()
# 分割后的句子
output_texts: List[str] = []
# 当前句子,在遇到句子边界预测之前逐步构建
current_chars: List[str] = []
# 遍历输出,忽略第一个(BOS)和最后一个(EOS)预测和 token
for token_idx in range(1, len(input_ids) - 1):
token = tokenizer.IdToPiece(input_ids[token_idx])
# 简单的 SP 解码
if token.startswith(\"▁\") and current_chars:
current_chars.append(\" \")
# token 级别的预测
pre_label = pre_labels[pre_preds[token_idx]]
post_label = post_labels[post_preds[token_idx]]
# 如果我们预测到"前置标点",则在此 token 之前插入
if pre_label != null_token:
current_chars.append(pre_label)
# 遍历每个字符。跳过 SP 的空格 token,
char_start = 1 if token.startswith(\"▁\") else 0
for token_char_idx, char in enumerate(token[char_start:], start=char_start):
# 如果此字符应大写,则应用大写
if cap_preds[token_idx][token_char_idx]:
char = char.upper()
# 追加字符
current_chars.append(char)
# 如果这是缩写,则在每个字符后添加句点(p.m.、a.m. 等)
if post_label == acronym_token:
current_chars.append(\".\")
# 可能此 subtoken 以标点符号结尾
if post_label != null_token and post_label != acronym_token:
current_chars.append(post_label)
# 如果此 token 是句子边界,则完成当前句子并重置
if sbd_preds[token_idx]:
output_texts.append(\"\".join(current_chars))
current_chars.clear()
# 如果最后一个 token 未被分类为句子边界,可能需要输出最后的句子
if current_chars:
output_texts.append(\"\".join(current_chars))
# 格式化输出
print(f\"Input: {input_text}\")
print(\"Outputs:\")
for text in output_texts:
print(f\"\t{text}\")
预期输出:
Input: hola mundo cómo estás estamos bajo el sol y hace mucho calor santa coloma abre los huertos urbanos a las escuelas de la ciudad
Outputs:
Hola mundo, ¿cómo estás?
Estamos bajo el sol y hace mucho calor.
Santa Coloma abre los huertos urbanos a las escuelas de la ciudad.
</details>
模型架构
该模型实现了以下计算图,允许在每种语言中进行标点符号、正确大小写和句号预测,而无需特定语言的行为:

<details>
<summary>点击查看计算图说明</summary>
我们首先对文本进行分词,并使用 XLM-Roberta 进行编码,这是该图的预训练部分。
然后我们预测每个 subtoken 前后的标点符号。在每个 token 之前进行预测可以实现西班牙语的反问号。在每个 token 之后进行预测可以实现所有其他标点符号,包括连续书写语言和缩写中的标点符号。
我们使用嵌入来表示预测的标点 token,以将标点信息传递给句子边界检测头。这允许正确预测句号,因为某些标点 token(句号、问号等)与句子边界有很强的相关性。
然后我们将句号预测向右移动一位,以告知正确大小写检测头每个新句子的开始位置。这很重要,因为正确大小写与句子边界有很强的相关性。
对于正确大小写,我们为每个 subtoken 预测 N 个结果,其中 N 是该 subtoken 中的字符数。实际上,N 是最大 subtoken 长度,多余的预测将被忽略。本质上,正确大小写被建模为一个多标签问题。这允许对任意字符进行大写,例如 "NATO"、"MacDonald"、"mRNA" 等。
将所有这些预测应用到输入文本中,我们可以对任何语言进行标点添加、正确大小写和句子分割。
</details>
分词器
<details>
<summary>点击查看 XLM-Roberta 分词器是如何被"修复"的</summary>
与 FairSeq 使用的 hacky 封装器以及 HuggingFace 奇怪地移植(而非修复)的方式不同,xlm-roberta SentencePiece 模型被调整为正确编码文本。根据 HF 的注释,
# 原始 fairseq 词汇表和 spm 词汇表必须"对齐":
# Vocab | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9
# -------- | ------- | ------- | ------ | ------- | --- | --- | --- | ----- | ----- | ----
# fairseq | '<s>' | '<pad>' | '</s>' | '<unk>' | ',' | '.' | '▁' | 's' | '▁de' | '-'
# spm | '<unk>' | '<s>' | '</s>' | ',' | '.' | '▁' | 's' | '▁de' | '-' | '▁a'
SP 模型通过以下代码片段进行了修复 (SentencePiece 专家,如果有问题请告诉我):
from sentencepiece import SentencePieceProcessor
from sentencepiece.sentencepiece_model_pb2 import ModelProto
m = ModelProto()
m.ParseFromString(open(\"/path/to/xlmroberta/sentencepiece.bpe.model\", \"rb\").read())
pieces = list(m.pieces)
pieces = (
[
ModelProto.SentencePiece(piece=\"<s>\", type=ModelProto.SentencePiece.Type.CONTROL),
ModelProto.SentencePiece(piece=\"<pad>\", type=ModelProto.SentencePiece.Type.CONTROL),
ModelProto.SentencePiece(piece=\"</s>\", type=ModelProto.SentencePiece.Type.CONTROL),
ModelProto.SentencePiece(piece=\"<unk>\", type=ModelProto.SentencePiece.Type.UNKNOWN),
]
+ pieces[3:]
+ [ModelProto.SentencePiece(piece=\"<mask>\", type=ModelProto.SentencePiece.Type.USER_DEFINED)]
)
del m.pieces[:]
m.pieces.extend(pieces)
with open(\"/path/to/new/sp.model\", \"wb\") as f:
f.write(m.SerializeToString())
现在我们可以仅使用 SP 模型而无需封装器。
</details>
后置标点 Token
该模型在每个 subtoken 之后预测以下标点 token 集合:
| Token | 描述 | 相关语言 |
|---|---|---|
| <NULL> | 无标点 | 所有 |
| <ACRONYM> | 此子词中的每个字符后跟一个句点 | 主要是英语,部分欧洲语言 |
| . | 拉丁文句号 | 多种 |
| , | 拉丁文逗号 | 多种 |
| ? | 拉丁文问号 | 多种 |
| ? | 全角问号 | 中文、日文 |
| , | 全角逗号 | 中文、日文 |
| 。 | 全角句号 | 中文、日文 |
| 、 | 表意逗号 | 中文、日文 |
| ・ | 中点 | 日文 |
| । | Danda | 印地语、孟加拉语、奥里亚语 |
| ؟ | 阿拉伯语问号 | 阿拉伯语 |
| ; | 希腊语问号 | 希腊语 |
| ። | 埃塞俄比亚语句号 | 阿姆哈拉语 |
| ፣ | 埃塞俄比亚语逗号 | 阿姆哈拉语 |
| ፧ | 埃塞俄比亚语问号 | 阿姆哈拉语 |
前置标点 Token
该模型在每个子词之前预测以下标点 token 集合:
| Token | 描述 | 相关语言 |
|---|---|---|
| <NULL> | 无标点 | 所有 |
| ¿ | 反问号 | 西班牙语 |
训练详情
该模型在 NeMo 框架中于 A100 上训练了约 7 小时。
您可以在 tensorboard.dev 上查看 tensorboard 日志。
该模型使用来自 WMT 的 News Crawl 数据进行训练。 每种语言使用 100 万行文本,少数低资源语言可能使用较少。 语言的选择基于 News Crawl 语料库是否包含足够多作者认为可靠质量的数据。
局限性
该模型是在新闻数据上训练的,在对话或非正式数据上可能表现不佳。
该模型不太可能达到生产质量。 它"仅"用每种语言 100 万行数据进行训练,且由于网络抓取的新闻数据的性质,开发集可能存在噪声。
该模型过度预测西班牙语问号,尤其是反问号 ¿(见下文指标)。
由于 ¿ 是一个稀有 token,尤其是在 47 语言模型的背景下,西班牙语问句被过采样,方法是从未使用的额外训练数据中选择更多这类句子。然而,这似乎"过度纠正"了问题,导致预测了许多西班牙语问号。
该模型也可能过度预测逗号。
如果您发现此处未提及的任何一般局限性,请告诉我,以便在下一次微调中解决所有局限性。
评估
在这些指标中,请记住
-
数据存在噪声
-
句子边界和正确大小写以预测的标点为条件,这是最困难的任务,有时会出错。 当以参考标点为条件时,大多数语言的正确大小写和 SBD 实际上达到 100%。
-
标点可能是主观的。例如,
Hola mundo, ¿cómo estás?或者
Hola mundo. ¿Cómo estás?当句子更长、更实际时,这些歧义比比皆是,并影响所有三项分析。
测试数据和示例生成
每个测试示例使用以下过程生成:
- 连接 11 个随机句子(测试集中每个句子 1 + 10 个)
- 将连接后的句子转换为小写
- 删除所有标点
目标在我们将字母转换为小写并删除标点时生成。
数据是 News Crawl 的保留部分,已去重。
每种语言使用 3,000 行数据,生成 3,000 个包含 11 个句子的唯一示例。
我们生成 3,000 个示例,其中示例 i 以句子 i 开头,后跟从 3,000 句测试集中随机选择的 10 个句子。
为了测量正确大小写和句子边界检测,使用参考标点 token 进行条件化(见上图)。如果我们改用预测的标点,则错误的标点将导致正确大小写和 SBD 目标无法正确对齐,这些指标将被人为降低。
选定语言评估报告
目前,下面显示了少数选定语言的指标。 鉴于收集和格式化 47 种语言的指标所需的工作量,我最终会添加更多。
展开以下任一选项卡以查看该语言的指标。
<details> <summary>英语</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.25 98.43 98.84 564908
<ACRONYM> (label_id: 1) 63.14 84.67 72.33 613
. (label_id: 2) 90.97 93.91 92.42 32040
, (label_id: 3) 73.95 84.32 78.79 24271
? (label_id: 4) 79.05 81.94 80.47 1041
? (label_id: 5) 0.00 0.00 0.00 0
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 0.00 0.00 0.00 0
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 97.60 97.60 97.60 622873
macro avg 81.27 88.65 84.57 622873
weighted avg 97.77 97.60 97.67 622873
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 99.72 99.85 99.78 2134956
UPPER (label_id: 1) 96.33 93.52 94.91 91996
-------------------
micro avg 99.59 99.59 99.59 2226952
macro avg 98.03 96.68 97.34 2226952
weighted avg 99.58 99.59 99.58 2226952
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.98 99.99 591540
FULLSTOP (label_id: 1) 99.61 99.89 99.75 34333
-------------------
micro avg 99.97 99.97 99.97 625873
macro avg 99.80 99.93 99.87 625873
weighted avg 99.97 99.97 99.97 625873
</details>
<details> <summary>西班牙语</summary>
punct_pre test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.94 99.89 99.92 636941
¿ (label_id: 1) 56.73 71.35 63.20 1288
-------------------
micro avg 99.83 99.83 99.83 638229
macro avg 78.34 85.62 81.56 638229
weighted avg 99.85 99.83 99.84 638229
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.19 98.41 98.80 578271
<ACRONYM> (label_id: 1) 30.10 56.36 39.24 55
. (label_id: 2) 91.92 93.12 92.52 30856
, (label_id: 3) 72.98 82.44 77.42 27761
? (label_id: 4) 52.77 71.85 60.85 1286
? (label_id: 5) 0.00 0.00 0.00 0
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 0.00 0.00 0.00 0
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 97.40 97.40 97.40 638229
macro avg 69.39 80.44 73.77 638229
weighted avg 97.60 97.40 97.48 638229
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 99.82 99.86 99.84 2324724
UPPER (label_id: 1) 95.92 94.70 95.30 79266
-------------------
micro avg 99.69 99.69 99.69 2403990
macro avg 97.87 97.28 97.57 2403990
weighted avg 99.69 99.69 99.69 2403990
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.96 99.98 607057
FULLSTOP (label_id: 1) 99.31 99.88 99.60 34172
-------------------
micro avg 99.96 99.96 99.96 641229
macro avg 99.65 99.92 99.79 641229
weighted avg 99.96 99.96 99.96 641229
</details>
<details> <summary>阿姆哈拉语</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.83 99.28 99.56 729664
<ACRONYM> (label_id: 1) 0.00 0.00 0.00 0
. (label_id: 2) 0.00 0.00 0.00 0
, (label_id: 3) 0.00 0.00 0.00 0
? (label_id: 4) 0.00 0.00 0.00 0
? (label_id: 5) 0.00 0.00 0.00 0
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 0.00 0.00 0.00 0
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 91.27 97.90 94.47 25341
፣ (label_id: 15) 61.93 82.11 70.60 5818
፧ (label_id: 16) 67.41 81.73 73.89 1177
-------------------
micro avg 99.08 99.08 99.08 762000
macro avg 80.11 90.26 84.63 762000
weighted avg 99.21 99.08 99.13 762000
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 98.40 98.03 98.21 1064
UPPER (label_id: 1) 71.23 75.36 73.24 69
-------------------
micro avg 96.65 96.65 96.65 1133
macro avg 84.81 86.69 85.73 1133
weighted avg 96.74 96.65 96.69 1133
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.85 99.92 743158
FULLSTOP (label_id: 1) 95.20 99.62 97.36 21842
-------------------
micro avg 99.85 99.85 99.85 765000
macro avg 97.59 99.74 98.64 765000
weighted avg 99.85 99.85 99.85 765000
</details>
<details> <summary>中文</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.53 97.31 98.41 435611
<ACRONYM> (label_id: 1) 0.00 0.00 0.00 0
. (label_id: 2) 0.00 0.00 0.00 0
, (label_id: 3) 0.00 0.00 0.00 0
? (label_id: 4) 0.00 0.00 0.00 0
? (label_id: 5) 81.85 87.31 84.49 1513
, (label_id: 6) 74.08 93.67 82.73 35921
。 (label_id: 7) 96.51 96.93 96.72 32097
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 97.00 97.00 97.00 505142
macro avg 87.99 93.81 90.59 505142
weighted avg 97.48 97.00 97.15 505142
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 94.89 94.98 94.94 2951
UPPER (label_id: 1) 81.34 81.03 81.18 796
-------------------
micro avg 92.02 92.02 92.02 3747
macro avg 88.11 88.01 88.06 3747
weighted avg 92.01 92.02 92.01 3747
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.97 99.98 473642
FULLSTOP (label_id: 1) 99.55 99.90 99.72 34500
-------------------
micro avg 99.96 99.96 99.96 508142
macro avg 99.77 99.93 99.85 508142
weighted avg 99.96 99.96 99.96 508142
</details>
<details> <summary>日语</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.34 95.90 97.59 406341
<ACRONYM> (label_id: 1) 0.00 0.00 0.00 0
. (label_id: 2) 0.00 0.00 0.00 0
, (label_id: 3) 0.00 0.00 0.00 0
? (label_id: 4) 0.00 0.00 0.00 0
? (label_id: 5) 70.55 73.56 72.02 1456
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 94.38 96.95 95.65 32537
、 (label_id: 8) 54.28 87.62 67.03 18610
・ (label_id: 9) 28.18 71.64 40.45 1100
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 95.51 95.51 95.51 460044
macro avg 69.35 85.13 74.55 460044
weighted avg 96.91 95.51 96.00 460044
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 92.33 94.03 93.18 4174
UPPER (label_id: 1) 83.51 79.46 81.43 1587
-------------------
micro avg 90.02 90.02 90.02 5761
macro avg 87.92 86.75 87.30 5761
weighted avg 89.90 90.02 89.94 5761
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.92 99.96 428544
FULLSTOP (label_id: 1) 99.07 99.87 99.47 34500
-------------------
micro avg 99.92 99.92 99.92 463044
macro avg 99.53 99.90 99.71 463044
weighted avg 99.92 99.92 99.92 463044
</details>
<details> <summary>印地语</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.75 99.44 99.59 560358
<ACRONYM> (label_id: 1) 0.00 0.00 0.00 0
. (label_id: 2) 0.00 0.00 0.00 0
, (label_id: 3) 69.55 78.48 73.75 8084
? (label_id: 4) 63.30 87.07 73.31 317
? (label_id: 5) 0.00 0.00 0.00 0
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 0.00 0.00 0.00 0
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 96.92 98.66 97.78 32118
؟ (label_id: 11) 0.00 0.00 0.00 0
، (label_id: 12) 0.00 0.00 0.00 0
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 99.11 99.11 99.11 600877
macro avg 82.38 90.91 86.11 600877
weighted avg 99.17 99.11 99.13 600877
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 97.19 96.72 96.95 2466
UPPER (label_id: 1) 89.14 90.60 89.86 734
-------------------
micro avg 95.31 95.31 95.31 3200
macro avg 93.17 93.66 93.41 3200
weighted avg 95.34 95.31 95.33 3200
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 100.00 99.99 99.99 569472
FULLSTOP (label_id: 1) 99.82 99.99 99.91 34405
-------------------
micro avg 99.99 99.99 99.99 603877
macro avg 99.91 99.99 99.95 603877
weighted avg 99.99 99.99 99.99 603877
</details>
<details> <summary>阿拉伯语</summary>
punct_post test report:
label precision recall f1 support
<NULL> (label_id: 0) 99.30 96.94 98.10 688043
<ACRONYM> (label_id: 1) 93.33 77.78 84.85 18
. (label_id: 2) 93.31 93.78 93.54 28175
, (label_id: 3) 0.00 0.00 0.00 0
? (label_id: 4) 0.00 0.00 0.00 0
? (label_id: 5) 0.00 0.00 0.00 0
, (label_id: 6) 0.00 0.00 0.00 0
。 (label_id: 7) 0.00 0.00 0.00 0
、 (label_id: 8) 0.00 0.00 0.00 0
・ (label_id: 9) 0.00 0.00 0.00 0
। (label_id: 10) 0.00 0.00 0.00 0
؟ (label_id: 11) 65.93 82.79 73.40 860
، (label_id: 12) 44.89 79.20 57.30 20941
; (label_id: 13) 0.00 0.00 0.00 0
። (label_id: 14) 0.00 0.00 0.00 0
፣ (label_id: 15) 0.00 0.00 0.00 0
፧ (label_id: 16) 0.00 0.00 0.00 0
-------------------
micro avg 96.29 96.29 96.29 738037
macro avg 79.35 86.10 81.44 738037
weighted avg 97.49 96.29 96.74 738037
cap test report:
label precision recall f1 support
LOWER (label_id: 0) 97.10 99.49 98.28 4137
UPPER (label_id: 1) 98.71 92.89 95.71 1729
-------------------
micro avg 97.55 97.55 97.55 5866
macro avg 97.90 96.19 96.99 5866
weighted avg 97.57 97.55 97.52 5866
seg test report:
label precision recall f1 support
NOSTOP (label_id: 0) 99.99 99.97 99.98 710456
FULLSTOP (label_id: 1) 99.39 99.85 99.62 30581
-------------------
micro avg 99.97 99.97 99.97 741037
macro avg 99.69 99.91 99.80 741037
weighted avg 99.97 99.97 99.97 741037
</details>
附加内容
缩写、缩略语和双大写词
本节简要演示模型在以下情况下的行为:
- 缩写:"NATO"
- 假缩写:用 "NHTG" 代替 "NATO"
- 可能是缩写或专有名词的歧义词:"Tuny"
- 双大写词:"McDavid"
- 首字母缩略词:"p.m."
<details open>
<summary>缩写等输入</summary>
from typing import List
from punctuators.models import PunctCapSegModelONNX
m: PunctCapSegModelONNX = PunctCapSegModelONNX.from_pretrained(
\"1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase\"
)
input_texts = [
\"the us is a nato member as a nato member the country enjoys security guarantees notably article 5\",
\"the us is a nhtg member as a nhtg member the country enjoys security guarantees notably article 5\",
\"the us is a tuny member as a tuny member the country enjoys security guarantees notably article 5\",
\"connor andrew mcdavid is a canadian professional ice hockey centre and captain of the edmonton oilers of the national hockey league the oilers selected him first overall in the 2015 nhl entry draft mcdavid spent his childhood playing ice hockey against older children\",
\"please rsvp for the party asap preferably before 8 pm tonight\",
]
results: List[List[str]] = m.infer(
texts=input_texts, apply_sbd=True,
)
for input_text, output_texts in zip(input_texts, results):
print(f\"Input: {input_text}\")
print(f\"Outputs:\")
for text in output_texts:
print(f\"\t{text}\")
print()
</details>
<details open>
<summary>预期输出</summary>
Input: the us is a nato member as a nato member the country enjoys security guarantees notably article 5
Outputs:
The U.S. is a NATO member.
As a NATO member, the country enjoys security guarantees, notably Article 5.
Input: the us is a nhtg member as a nhtg member the country enjoys security guarantees notably article 5
Outputs:
The U.S. is a NHTG member.
As a NHTG member, the country enjoys security guarantees, notably Article 5.
Input: the us is a tuny member as a tuny member the country enjoys security guarantees notably article 5
Outputs:
The U.S. is a Tuny member.
As a Tuny member, the country enjoys security guarantees, notably Article 5.
Input: connor andrew mcdavid is a canadian professional ice hockey centre and captain of the edmonton oilers of the national hockey league the oilers selected him first overall in the 2015 nhl entry draft mcdavid spent his childhood playing ice hockey against older children
Outputs:
Connor Andrew McDavid is a Canadian professional ice hockey centre and captain of the Edmonton Oilers of the National Hockey League.
The Oilers selected him first overall in the 2015 NHL entry draft.
McDavid spent his childhood playing ice hockey against older children.
Input: please rsvp for the party asap preferably before 8 pm tonight
Outputs:
Please RSVP for the party ASAP, preferably before 8 p.m. tonight.
</details>
Salama1429/xlm-roberta_punctuation_fullstop_truecase
作者 Salama1429
创建时间: 2025-01-30 14:07:37+00:00
更新时间: 2025-02-02 11:21:52+00:00
在 Hugging Face 上查看