Compatibility and performance results, a same-document comparison, and a reproducible deployment procedure
Redacted public edition · 2026-09-27 · Kunshan Z100 / gfx906 · Transformers/PyTorch backend
PaddleOCR 3.7.0 completed end-to-end functional verification on Kunshan Z100 through the Transformers/PyTorch backend. The deployment uses a user-space glibc 2.28 loader, GCC 12.2, DTK 26.04, and a Hygon-adapted Torch 2.7.1 build. It does not depend on a PaddlePaddle DCU wheel.
| Layer | Configuration | Notes |
|---|---|---|
| Hardware | Hygon Z100 / gfx906 / Device 66a1 | One DCU requested for the job |
| Operating system | CentOS 7.6 / system glibc 2.17 | User-space glibc 2.28 handles the wheel's ABI requirements |
| Toolchain | GCC 12.2 + DTK 26.04 | DTK supplies HIP and math libraries; GCC supplies a newer libstdc++ |
| Python | 3.10.14 | Reuses the existing base interpreter without modifying it |
| PyTorch | 2.7.1+das.opt1.dtk2604 / HIP 6.3.26093 | Hygon-adapted build; a generic PyPI PyTorch wheel is not a substitute |
| Transformers | 5.17.0 / Hub 1.5.0 / Tokenizers 0.23.1 | Versions installed in the separate PaddleOCR environment |
| Application | PaddleOCR 3.7.0 / PaddleX 3.7.2 | Separate package tree; MinerU's virtual environment is unchanged |
The available Kunshan modules are limited mainly to Python 3.7 and PyTorch 1.10. PaddleOCR 3.7 and PaddleX 3.7 require Python 3.8 syntax and a modern Transformers API, so installing them directly into Python 3.7 fails on syntax and typing/importlib APIs.
The MinerU work established a more suitable user-space runtime: Python 3.10, PyTorch 2.7.1, DTK 26.04, and glibc 2.28. This test reuses its PyTorch and runtime libraries read-only, leaves MinerU unchanged, and installs PaddleOCR/PaddleX in a separate directory.
Compute nodes have no external network access. The model is therefore downloaded from BOS to a new cache on the login node. On compute nodes, HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 prevent runtime DNS timeouts.
| Probe | Result |
|---|---|
torch.cuda.is_available() | True |
| Device name / architecture | Device 66a1 / gfx906:sramecc-:xnack- |
| GPU matrix multiplication | Completed and synchronized on cuda:0 |
| PaddleOCR/PaddleX import | 3.7.0 / 3.7.2 |
| PP-OCRv5 mobile det | Model loaded; detection boxes returned |
| PP-OCRv5 mobile rec | Model loaded; recognized 16 text lines |
| Full OCR | PADDLEOCR_TRANSFORMERS_E2E_OK, exit code 0 |
| Metric | Measured value |
|---|---|
| Cold model load | 67.6–80.1 s |
| Initial warm-up | Approximately 14–21.5 s |
| Mean over five unique inputs | 0.2546 s/image |
| P50 / P95 | 0.2430 / 0.3013 s |
| Steady-state throughput | Approximately 3.93 images/s |
| Recognized lines per image | 16 |
| GPU memory during inference | 0.212 GB allocated / 0.281 GB reserved |
| Stage | Measured value |
|---|---|
| PDF rendering (pypdfium2, scale=2.0) | 0.7254 s |
| Model load | 80.0806 s |
| OCR warmup | 21.5097 s |
| Steady-state OCR for 23 pages | 46.9364 s |
| Steady-state mean | 2.0404 s/page |
| P50 / P95 | 1.9419 / 3.0614 s/page |
| Steady-state throughput | 0.49 pages/s |
| Cold start, including warm-up | Approximately 149.3 s |
Both systems processed the same 23-page, approximately 0.58 MB file, Mooncake-v3.pdf.
MinerU figures are taken from its corresponding build report; PaddleOCR was measured again on that PDF for this report.
| Dimensions | MinerU pipeline | PaddleOCR Transformers |
|---|---|---|
| Cold start | 146.8 s | About 149.3 s |
| Steady-state runtime | 90.5 s / 23 pages | 46.94 s / 23 pages |
| Mean per page | 3.93 s/page | 2.04 s/page |
| Text characters | 81,072 Markdown characters | 81,550 OCR characters |
| Output structure | Headings, tables, images, formulas, middle.json, and layout PDF | Text boxes, text, and confidence scores |
| Formulas and tables | 17 inline formulas and 3 tables | No structured reconstruction |
| Images | 23 references / 26 files | No image extraction |
No character-level ground truth was prepared for this test, so character-count differences are not treated as accuracy measurements. PaddleOCR's mean confidence was 0.9636, yet spot checks of the same output still found errors such as:
Mooncacke
diversifi ed
work- loads
2O24
The MinerU report separately verifies heading hierarchy, author information, table values, formulas, and extracted images. For papers, contracts, forms, and PDF-to-Markdown workflows, structured output matters more; PaddleOCR is a better fit for batch plain-text recognition, targeted re-recognition, and low-latency OCR.
The system glibc 2.17 cannot load the newer Torch wheel directly. Production must run through the user-space glibc 2.28 loader, with GCC 12.2, DTK 26.04, dcc/gcvm, and rocm_smi in its runtime environment.
<PADDLEOCR_ROOT>/bin/dcu-python your_worker.py
Do not invoke the system python directly. Spawned child processes must not fall back to the unwrapped sys.executable; the environment's sitecustomize.py sets multiprocessing.set_executable().
torch.cuda.device_count() == 1
torch.cuda.get_device_properties(0).gcnArchName startswith("gfx906")
PADDLEOCR_TRANSFORMERS_E2E_OK
PADDLEOCR_EXIT_RC=0
summary.json exists and chars/lines > 0
The report package includes the following files:
| File | Purpose |
|---|---|
paddleocr_z100_launcher.sh | User mode glibc 2.28 + GCC/DTK loader template |
paddleocr_z100_benchmark.py | PDF rendering, model loading, warm-up, page-level OCR, GPU-memory use, and text statistics |
paddleocr_z100.slurm | Single-card Z100 Slurm job template |
README.md | Installation, model caching, offline operation, and deployment guidance |
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
export PADDLE_PDX_CACHE_HOME=<PADDLEOCR_ROOT>/cache
export PADDLE_PDX_MODEL_SOURCE=BOS
export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
<PADDLEOCR_ROOT>/bin/dcu-python \
paddleocr_z100_benchmark.py \
<INPUT_PDF> <OUTPUT_DIR>
| Acceptance item | Result |
|---|---|
| Independent environment | Passed; MinerU venv unchanged |
| Offline model cache | PP-OCRv5 mobile detection/recognition safetensors |
| Real Z100 | 1 device 66a1 / gfx906 |
| End-to-end OCR | 16 test lines recognized; exit code 0 |
| 23-page document | Completed; 1,513 lines and 81,550 characters |
| Performance statistics | summary.json generated |
The public report omits account, host, node, and job identifiers, as well as real absolute paths and internal service addresses. Replace the placeholders in the local launcher and Slurm templates when reproducing the test.
Accuracy:No OCR ground truth is available yet. Before production use, evaluate representative data at the character, line, or field level.
Performance:The attention and BLAS paths on gfx906 fall back to less optimized implementations. Rebenchmark after upgrading DTK or Torch.
Model:This test uses PP-OCRv5 mobile detection and recognition. The default PP-OCRv6 path requires additional model downloads, while compute nodes currently have no external network access.
Concurrency:Only one card and one worker were tested. A production deployment should use one process per card and first validate throughput and stability with 1, 2, and 4 cards.
Document parsing:For workflows that require tables, formulas, images, and Markdown structure, continue using MinerU and use PaddleOCR as an OCR component where appropriate.
Environment drift:If the loader, model cache, modules, Torch, or Transformers are moved or upgraded, repeat the reproduction steps in §9.