Reproducible benchmark

PDF-to-Word benchmark

This page publishes reproducible methods, data provenance, and results. ChuyenFile does not publish a quality score before the corpus and manual-review gates are met.

Corpus collection — not certified
Documents
100
Minimum gate: 30
Pages
100
Minimum gate: 100
CER
0.69
WER
1.7

Methodology

  1. Separate development and acceptance splits; never tune on the acceptance split.
  2. Measure CER/WER after Unicode NFC normalization; never substitute OCR confidence for accuracy.
  3. Check page count, dimensions, text nodes, tables, and renders; manually review complex layouts.
  4. Publish failures and weak document categories, not only averages.

Dataset provenance

Omni OCR Benchmark

Public OCR images with reference markdown, deterministically sampled and converted to one-page PDFs

License/use constraint: MIT

ChuyenFile controlled Vietnamese set

Vietnamese ground truth, tables, fonts and controlled scan degradation; published when its generator and artifacts are complete

License/use constraint: project-generated test data

Machine-readable artifacts

The current JSON report explicitly reports an uncertified state instead of inventing metrics.

pdf-to-word.json