Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
finepdfs_ocr_training_v1_sample – Dataset by mananbruce | AlphaNeural AI
You can deploy this model and start earning money today!
mananbruce
/
finepdfs_ocr_training_v1_sample
like
0
image-to-text
document-question-answering
zh
other
n<1K
parquet
image
text
document
datasets
dask
polars
mlcroissant
us
ocr
document-understanding
chinese
pdf
parquet
sample
Views
No views yet
Model card
Files and Versions
Community
API
FinePDFs OCR Training V1 Sample
这是 FinePDFs 中文 PDF OCR 训练数据 HuggingFace 发布格式的小样例。每条样本对应一个 PDF 页面图像,包含原图、MinerU doctag 伪标签、GLM/Paddle doctag 参考结果、一致性评分指标和训练目标分类。
规模
分区 样本数
text_no_table 8
text_table_region_only 8
table_expert_only 8
text_table_region_and_table_expert 8
total unique rows 32
总大小约 11.14 MiB。
构建流程
数据集按下面流程构建:
从 FinePDFs 中文 PDF 数据中下载 PDF,并过滤非简体中文材料。 将 PDF 拆分为单页图像,记录页面来源、尺寸、文件哈希等元信息。 对页面图像进行去重、图片质量过滤和 embedding… See the full description on the dataset page:
https://huggingface.co/datasets/mananbruce/finepdfs_ocr_training_v1_sample
.