opensource_community

PaddleOCR on Hygon Z100

Compatibility and performance results, a same-document comparison, and a reproducible deployment procedure

Redacted public edition · 2026-09-27 · Kunshan Z100 / gfx906 · Transformers/PyTorch backend

PaddleOCR 3.7.0PaddleX 3.7.2 Torch 2.7.1+DTK2604Python 3.10.14 Single card Z100

1. Final conclusion

PaddleOCR 3.7.0 completed end-to-end functional verification on Kunshan Z100 through the Transformers/PyTorch backend. The deployment uses a user-space glibc 2.28 loader, GCC 12.2, DTK 26.04, and a Hygon-adapted Torch 2.7.1 build. It does not depend on a PaddlePaddle DCU wheel.

PassedDevice enumeration: 1 × Device 66a1 / gfx906
0.2546 sSingle image stable end-to-end OCR (v2.1 test image)
0.49 pages/s23 pages Mooncake PDF, steady state OCR
0End-to-end failures
PaddleOCR is a high-throughput text detection and recognition engine; MinerU is a document-parsing pipeline. Character counts alone are not a meaningful basis for deciding whether one can replace the other.

2. Test environment and software configuration

LayerConfigurationNotes
HardwareHygon Z100 / gfx906 / Device 66a1One DCU requested for the job
Operating systemCentOS 7.6 / system glibc 2.17User-space glibc 2.28 handles the wheel's ABI requirements
ToolchainGCC 12.2 + DTK 26.04DTK supplies HIP and math libraries; GCC supplies a newer libstdc++
Python3.10.14Reuses the existing base interpreter without modifying it
PyTorch2.7.1+das.opt1.dtk2604 / HIP 6.3.26093Hygon-adapted build; a generic PyPI PyTorch wheel is not a substitute
Transformers5.17.0 / Hub 1.5.0 / Tokenizers 0.23.1Versions installed in the separate PaddleOCR environment
ApplicationPaddleOCR 3.7.0 / PaddleX 3.7.2Separate package tree; MinerU's virtual environment is unchanged

3. Inference backend selection

The available Kunshan modules are limited mainly to Python 3.7 and PyTorch 1.10. PaddleOCR 3.7 and PaddleX 3.7 require Python 3.8 syntax and a modern Transformers API, so installing them directly into Python 3.7 fails on syntax and typing/importlib APIs.

The MinerU work established a more suitable user-space runtime: Python 3.10, PyTorch 2.7.1, DTK 26.04, and glibc 2.28. This test reuses its PyTorch and runtime libraries read-only, leaves MinerU unchanged, and installs PaddleOCR/PaddleX in a separate directory.

Compute nodes have no external network access. The model is therefore downloaded from BOS to a new cache on the login node. On compute nodes, HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 prevent runtime DNS timeouts.

4. Compatibility verification

ProbeResult
torch.cuda.is_available()True
Device name / architectureDevice 66a1 / gfx906:sramecc-:xnack-
GPU matrix multiplicationCompleted and synchronized on cuda:0
PaddleOCR/PaddleX import3.7.0 / 3.7.2
PP-OCRv5 mobile detModel loaded; detection boxes returned
PP-OCRv5 mobile recModel loaded; recognized 16 text lines
Full OCRPADDLEOCR_TRANSFORMERS_E2E_OK, exit code 0
Architecture limitation:gfx906 has no matching HIPBLASLt Tensile kernel, and FlashAttention is unavailable, so Transformers falls back to PyTorch's native math attention. Functional tests pass, but this is not the hardware's maximum achievable performance.

5. Performance results

5.1 Single-image steady state

MetricMeasured value
Cold model load67.6–80.1 s
Initial warm-upApproximately 14–21.5 s
Mean over five unique inputs0.2546 s/image
P50 / P950.2430 / 0.3013 s
Steady-state throughputApproximately 3.93 images/s
Recognized lines per image16
GPU memory during inference0.212 GB allocated / 0.281 GB reserved

23-page Mooncake-v3 PDF

StageMeasured value
PDF rendering (pypdfium2, scale=2.0)0.7254 s
Model load80.0806 s
OCR warmup21.5097 s
Steady-state OCR for 23 pages46.9364 s
Steady-state mean2.0404 s/page
P50 / P951.9419 / 3.0614 s/page
Steady-state throughput0.49 pages/s
Cold start, including warm-upApproximately 149.3 s

6. Same-document comparison with MinerU

Both systems processed the same 23-page, approximately 0.58 MB file, Mooncake-v3.pdf. MinerU figures are taken from its corresponding build report; PaddleOCR was measured again on that PDF for this report.

DimensionsMinerU pipelinePaddleOCR Transformers
Cold start146.8 sAbout 149.3 s
Steady-state runtime90.5 s / 23 pages46.94 s / 23 pages
Mean per page3.93 s/page2.04 s/page
Text characters81,072 Markdown characters81,550 OCR characters
Output structureHeadings, tables, images, formulas, middle.json, and layout PDFText boxes, text, and confidence scores
Formulas and tables17 inline formulas and 3 tablesNo structured reconstruction
Images23 references / 26 filesNo image extraction
Conclusion:PaddleOCR's steady-state OCR was about 1.9× faster. MinerU, however, is a document-parsing pipeline that also handles layout, tables, formulas, and images. Similar character counts do not establish equivalent accuracy, and they do not make PaddleOCR a substitute for MinerU.

7. Quality and accuracy limits

No character-level ground truth was prepared for this test, so character-count differences are not treated as accuracy measurements. PaddleOCR's mean confidence was 0.9636, yet spot checks of the same output still found errors such as:

Mooncacke
diversifi ed
work- loads
2O24

The MinerU report separately verifies heading hierarchy, author information, table values, formulas, and extracted images. For papers, contracts, forms, and PDF-to-Markdown workflows, structured output matters more; PaddleOCR is a better fit for batch plain-text recognition, targeted re-recognition, and low-latency OCR.

8. Deployment plan

8.1 Launcher

The system glibc 2.17 cannot load the newer Torch wheel directly. Production must run through the user-space glibc 2.28 loader, with GCC 12.2, DTK 26.04, dcc/gcvm, and rocm_smi in its runtime environment.

<PADDLEOCR_ROOT>/bin/dcu-python your_worker.py

Do not invoke the system python directly. Spawned child processes must not fall back to the unwrapped sys.executable; the environment's sitecustomize.py sets multiprocessing.set_executable().

8.2 Resident worker

8.3 Production acceptance

torch.cuda.device_count() == 1
torch.cuda.get_device_properties(0).gcnArchName startswith("gfx906")
PADDLEOCR_TRANSFORMERS_E2E_OK
PADDLEOCR_EXIT_RC=0
summary.json exists and chars/lines > 0

9. Reproduction entry points

The report package includes the following files:

FilePurpose
paddleocr_z100_launcher.shUser mode glibc 2.28 + GCC/DTK loader template
paddleocr_z100_benchmark.pyPDF rendering, model loading, warm-up, page-level OCR, GPU-memory use, and text statistics
paddleocr_z100.slurmSingle-card Z100 Slurm job template
README.mdInstallation, model caching, offline operation, and deployment guidance
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
export PADDLE_PDX_CACHE_HOME=<PADDLEOCR_ROOT>/cache
export PADDLE_PDX_MODEL_SOURCE=BOS
export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1

<PADDLEOCR_ROOT>/bin/dcu-python \
  paddleocr_z100_benchmark.py \
  <INPUT_PDF> <OUTPUT_DIR>

10. Test results and acceptance

Acceptance itemResult
Independent environmentPassed; MinerU venv unchanged
Offline model cachePP-OCRv5 mobile detection/recognition safetensors
Real Z1001 device 66a1 / gfx906
End-to-end OCR16 test lines recognized; exit code 0
23-page documentCompleted; 1,513 lines and 81,550 characters
Performance statisticssummary.json generated

The public report omits account, host, node, and job identifiers, as well as real absolute paths and internal service addresses. Replace the placeholders in the local launcher and Slurm templates when reproducing the test.

11. Known limitations and follow-up suggestions

Accuracy:No OCR ground truth is available yet. Before production use, evaluate representative data at the character, line, or field level.

Performance:The attention and BLAS paths on gfx906 fall back to less optimized implementations. Rebenchmark after upgrading DTK or Torch.

Model:This test uses PP-OCRv5 mobile detection and recognition. The default PP-OCRv6 path requires additional model downloads, while compute nodes currently have no external network access.

Concurrency:Only one card and one worker were tested. A production deployment should use one process per card and first validate throughput and stability with 1, 2, and 4 cards.

Document parsing:For workflows that require tables, formulas, images, and Markdown structure, continue using MinerU and use PaddleOCR as an OCR component where appropriate.

Environment drift:If the loader, model cache, modules, Torch, or Transformers are moved or upgraded, repeat the reproduction steps in §9.