opensource_community
Compatibility and performance report · Redacted public edition

MinerU on Hygon Z100

This report establishes MinerU 3.4.4 as the production baseline and evaluates MinerU 4.0.7 as an experimental option. It records the deployment configuration and supporting tests; the companion archive contains detailed failure logs and reproduction commands.

Hardware Z100 · gfx906 · 16 GB × 4 System CentOS 7.6 · glibc 2.17 Toolchain GCC 12.2 · DTK 26.04 Version MinerU 3.4.4 / 4.0.7 Scheduling Slurm
01

Main conclusions

QuestionConclusionStatus
Production processing3.4.4 pipeline + resident Z100 workerBaseline
Fast preview and text extraction4.0.7 FlashSuitable
Higher-quality 4.x local parsing4.0.7 Basic: ONNX on CPU or a resident Torch/DCU workerSuitable
Standard / AdvancedTransformers VLM isolation patch has passed functional verification and is for experimental use onlyExperimental
Batch submission interfaceSlurm-native mineru-drop, supports 3.x / 4.x adapterVerified
955/h
3.4.4 Four-card service pool throughput
12 s
3.4.4 pipeline, 14-page warm run
2.64 s
4.0.7 Basic/Torch, two-page warm run
4 / 4
4.0.7 functional checks passed
Summary

MinerU 4.0.7 has passed functional checks, but the available production evidence does not justify replacing 3.4.4. Keep the production and experimental environments separate; do not upgrade the production environment in place.

02

Test scope and method

Results from different MinerU modes, warm and Cold run runs, and document lengths are not combined into a single speed ranking.

EvidencePurposeRestrictions
Same PDF, page range, and resource allocationCold comparison of the 3.4.4 pipeline and 4.0.7 BasicOutput protocols differ
Repeated processing in one processMeasure warm throughput with the model kept residentExcludes Slurm queue time
Service pool concurrent stress testTest production throughput, back pressure and fault recoveryOnly 3.4.4 completed long-term verification
Single-page VLM smoke testConfirm that Standard / Advanced can runCannot be extrapolated to long-document throughput
  • Cold run: New process, including import, model loading and first execution.
  • Warm: Reuse the loaded model in the same process.
  • Flash and Basic/pipeline provide different output quality, so their throughput is not compared as if the results were quality-equivalent.
03

Test platform and runtime environment

CapabilitiesActual measurementMeaning
FP32 GEMMAbout 9.1 TFLOPSStable across retests
FP16 GEMMAbout 16.1 TFLOPSapproximately 1.8× FP32
BF16 GEMMAbout 5.8 TFLOPSNo hardware acceleration for this workload
Triton / torch.compileNot availablegfx906 and DTK runtime boundaries
vLLM HIP kernelPagedAttention can rundoes not establish compatibility of the complete vLLM service

3.1 Unified runtime entry point

CentOS 7's glibc 2.17 cannot directly load some wheels. Both environments reuse the same principle:

module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04

<GLIBC_2_28_LD_SO> \
  --library-path "<GLIBC_LIB>:<GCC_LIB>:<DTK_LIBS>:<SYSTEM_LIBS>" \
  <VENV>/bin/python "$@"

sitecustomize.py also directs multiprocessing spawn workers through the same launcher, preventing child processes from falling back to the system glibc.

04

MinerU 3.4.4 production baseline

The case for 3.4.4 rests on completed validation of output quality, steady-state throughput, backpressure, retries, and failure recovery—not simply on its being the older version.

ProjectResult
Backendpipeline
4-card service pool955 docs/hour, concurrently 12
Stability72 requests, 0 failed, 0 rejected
Warm document runAbout 12 seconds for 14 pages (0.86 s/page)
Fault recoveryUnder load, a killed worker is restarted automatically and its request is not lost
MemoryNo RSS increase across 72 requests
client │ ▼ Slurm Service Pool Gateway ── Backpressure / Least Load / Retry │ ├── Z100 worker 0 ├── Z100 worker 1 ├── Z100 worker 2 └── Z100 worker 3
Reason for retention

It is still the only path currently with complete production stress test evidence. Unless 4.0.7 completes equivalent corpus, warm service and failure experiments, the default production adapter will not be switched.

05

MinerU 4.0.7 experimental environment and validation

MinerU 4.0.7 uses a separate virtual environment, model directory, and configuration; it does not overwrite the 3.4.4 installation. In 4.x, tier specifies the quality level, while small_backend and vlm.engine select the underlying inference engines.

Old 3.x4.0.7 corresponds toimplements
Fast native analysisFlashNative text / Flash OCR
pipelineBasicONNX or Torch small model
hybrid-engineStandardsmall model + VLM
vlm-engineAdvancedHigher VLM inference investment

5.1 Experimental environment configuration

MinerU             4.0.7
DocVortex          0.5.1
ONNX Runtime       1.20.1
Torch              2.7.1 + DTK 26.04
torchvision        0.22.0 + DTK 26.04
Transformers       5.17.0
NumPy              1.26.4
OpenCV             4.11.0

The official Linux llama.cpp wheel requires glibc 2.34, while this user-space runtime provides glibc 2.28. Standard and Advanced therefore use isolated patches to expose the Transformers backends already present in the source tree.

06

Basic-mode performance optimization

The performance bottleneck is not the single-page operator, but "create a new Python + load all models for each task".

6.1 ONNX Runtime CPU thread configuration

intra-op threadCold 2 pagesWarm 2 pagesSuggestions
476.29 s10.43 sResource conservation
831.63 s6.60 sTotal throughput priority
1633.64 s6.13 sSingle worker delay priority

6.2 Torch on Z100

Status2 pages takechanges
Cold run198.12 sModel loading dominates
Warm run, SDPA math fallback2.86 sReuse a single loaded model
warm, eager formula attention2.64 sAbout 8% faster
has landed

mineru4-drop keeps one MinerUParser resident in each Slurm worker and disables formula SDPA on gfx906. When a 14-page document is split into 8- and 6-page chunks, the first chunk takes 148.8 seconds and the second takes 10.3 seconds.

07

3.4.4 and 4.0.7 controlled comparison

Both runs use the same document PDF (pages 1–2), 8 CPU cores, and 27 GB of memory. Each DCU run uses one Z100 card.

version/pathSlurm total time takenmodel stageJudgment
3.4.4 pipeline / CPU165 s36.95 sBaseline
4.0.7 Basic / ONNX CPU93 s37.15 sPeripheral startup is shorter
3.4.4 pipeline / Z100236 s60.02 sFaster Cold run DCU run
4.0.7 Basic / Torch Z100293 s71.97 sSlower on a cold run
  • CPU model stage is almost unchanged; 4.0.7 cold-start advantage comes from import, preprocessing and output chain.
  • DCU the 4.0.7 cold run is slower, but warm drops to 2.64 seconds / 2 pages.
  • The versions use different output protocols, so total runtime is an application-level comparison, not a kernel-only benchmark.
08

MinerU 4.0.7 functional verification

TierResources / BackendSampleResultPositioning
FlashCPU native text14 pages69 s operation, model stage 16.8 sPreview / Index
BasicONNX CPU / Torch Z1002-page and 14-page blocksValidated and optimizedLocal high quality
StandardTorch + Transformers VLM1 page Cold run160.9 sExperimental
AdvancedTransformers VLM1-page warm run24.3 sExperimental
Do not misread

Standard was measured on a a cold run, while Advanced reused the VLM loaded in the same process. These results are not directly comparable in quality or speed; they show only that all four tiers can parse a document in the target environment.

09

MinerU Drop batch-processing interface

native push file/directory │ ▼ Shared inbox ─ SHA256 / number of pages / chunking plan │ ▼ Slurm job array ─ One DCU per worker; model stays resident; tasks are claimed atomically │ ▼ CPU merge afterany ── Markdown / JSON / images / manifest │ ▼ status / retry / fetch

9.1 Adapter implementation

remote rootversionPurpose
mineru-drop3.4.4Stable production default
mineru4-drop4.0.7 Basic/TorchNew protocol and experimental batch processing
# 3.4.4
mineru-drop push <PDF_OR_DIRECTORY>

# 4.0.7
mineru-drop --remote-root mineru4-drop push <PDF_OR_DIRECTORY>

# Generic commands
mineru-drop [--remote-root ...] status --batch <BATCH_ID>
mineru-drop [--remote-root ...] retry  --batch <BATCH_ID>
mineru-drop [--remote-root ...] fetch  --batch <BATCH_ID>

Documents of up to 128 pages are processed as one batch; longer documents are split into 96-page chunks. For a small number of long documents, use one worker to avoid repeated model Cold run starts. Scale out to four workers for larger batches.

10

Deployment recommendations

10.1 Batch document processing

3.4.4 pipeline
has validated four-card throughput, stability, backpressure, and failure recovery.

10.2 Quick search and preview

4.0.7 Flash
requires no local large model and suits an initial scan or indexing pass.

10.3 4.x output workflow

4.0.7 Basic
CPU uses ONNX 8/16 threads; Z100 must be a resident model.

10.4 Evaluation of complex-page parsing

Standard / Advanced
Use only for manually selected small batches; exclude them from default automatic routing.

10.5 Recommended workflow

Native text; indexing only       → 4.0.7 Flash
Final Markdown / JSON required   → 3.4.4 pipeline
4.x schema required              → 4.0.7 Basic
Complex pages; manual review     → Standard / Advanced
11

Reproduction procedure and scope

11.1 Minimum reproduction requirements

  • The login node prepares dependencies and models; the computing node runs offline.
  • module must be loaded before venv.
  • The main process and the spawn child process must share the glibc 2.28 launcher.
  • Slurm scripts must pass real workload exit codes.
  • 4.0.7 ONNX Runtime 1.20.1 uses relabeled wheel and can only run under external glibc 2.28.

11.2 Known limitations

LimitationImpact
No FlashAttention on gfx906Torch formula/VLM attention fallback; the formula model has been switched to eager
mineru-llama-cpp requires glibc 2.34the official local Standard/Advanced engine cannot be installed
The vendor vLLM version predates MinerU 4.0.7Not used as the 4.x default VLM engine
Cross-block paragraphs and tablesThe manifest flags these cases for sampled review
Initial experiment: 2026-08-12 ~ 2026-08-13 · Dual version reconstruction: 2026-09-27 · Redacted public edition