MinerU on Hygon Z100
This report establishes MinerU 3.4.4 as the production baseline and evaluates MinerU 4.0.7 as an experimental option. It records the deployment configuration and supporting tests; the companion archive contains detailed failure logs and reproduction commands.
Main conclusions
| Question | Conclusion | Status |
|---|---|---|
| Production processing | 3.4.4 pipeline + resident Z100 worker | Baseline |
| Fast preview and text extraction | 4.0.7 Flash | Suitable |
| Higher-quality 4.x local parsing | 4.0.7 Basic: ONNX on CPU or a resident Torch/DCU worker | Suitable |
| Standard / Advanced | Transformers VLM isolation patch has passed functional verification and is for experimental use only | Experimental |
| Batch submission interface | Slurm-native mineru-drop, supports 3.x / 4.x adapter | Verified |
MinerU 4.0.7 has passed functional checks, but the available production evidence does not justify replacing 3.4.4. Keep the production and experimental environments separate; do not upgrade the production environment in place.
Test scope and method
Results from different MinerU modes, warm and Cold run runs, and document lengths are not combined into a single speed ranking.
| Evidence | Purpose | Restrictions |
|---|---|---|
| Same PDF, page range, and resource allocation | Cold comparison of the 3.4.4 pipeline and 4.0.7 Basic | Output protocols differ |
| Repeated processing in one process | Measure warm throughput with the model kept resident | Excludes Slurm queue time |
| Service pool concurrent stress test | Test production throughput, back pressure and fault recovery | Only 3.4.4 completed long-term verification |
| Single-page VLM smoke test | Confirm that Standard / Advanced can run | Cannot be extrapolated to long-document throughput |
- Cold run: New process, including import, model loading and first execution.
- Warm: Reuse the loaded model in the same process.
- Flash and Basic/pipeline provide different output quality, so their throughput is not compared as if the results were quality-equivalent.
Test platform and runtime environment
| Capabilities | Actual measurement | Meaning |
|---|---|---|
| FP32 GEMM | About 9.1 TFLOPS | Stable across retests |
| FP16 GEMM | About 16.1 TFLOPS | approximately 1.8× FP32 |
| BF16 GEMM | About 5.8 TFLOPS | No hardware acceleration for this workload |
| Triton / torch.compile | Not available | gfx906 and DTK runtime boundaries |
| vLLM HIP kernel | PagedAttention can run | does not establish compatibility of the complete vLLM service |
3.1 Unified runtime entry point
CentOS 7's glibc 2.17 cannot directly load some wheels. Both environments reuse the same principle:
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
<GLIBC_2_28_LD_SO> \
--library-path "<GLIBC_LIB>:<GCC_LIB>:<DTK_LIBS>:<SYSTEM_LIBS>" \
<VENV>/bin/python "$@"
sitecustomize.py also directs multiprocessing spawn workers through the same launcher, preventing child processes from falling back to the system glibc.
MinerU 3.4.4 production baseline
The case for 3.4.4 rests on completed validation of output quality, steady-state throughput, backpressure, retries, and failure recovery—not simply on its being the older version.
| Project | Result |
|---|---|
| Backend | pipeline |
| 4-card service pool | 955 docs/hour, concurrently 12 |
| Stability | 72 requests, 0 failed, 0 rejected |
| Warm document run | About 12 seconds for 14 pages (0.86 s/page) |
| Fault recovery | Under load, a killed worker is restarted automatically and its request is not lost |
| Memory | No RSS increase across 72 requests |
It is still the only path currently with complete production stress test evidence. Unless 4.0.7 completes equivalent corpus, warm service and failure experiments, the default production adapter will not be switched.
MinerU 4.0.7 experimental environment and validation
MinerU 4.0.7 uses a separate virtual environment, model directory, and configuration; it does not overwrite the 3.4.4 installation. In 4.x, tier specifies the quality level, while small_backend and vlm.engine select the underlying inference engines.
| Old 3.x | 4.0.7 corresponds to | implements |
|---|---|---|
| Fast native analysis | Flash | Native text / Flash OCR |
| pipeline | Basic | ONNX or Torch small model |
| hybrid-engine | Standard | small model + VLM |
| vlm-engine | Advanced | Higher VLM inference investment |
5.1 Experimental environment configuration
MinerU 4.0.7
DocVortex 0.5.1
ONNX Runtime 1.20.1
Torch 2.7.1 + DTK 26.04
torchvision 0.22.0 + DTK 26.04
Transformers 5.17.0
NumPy 1.26.4
OpenCV 4.11.0
The official Linux llama.cpp wheel requires glibc 2.34, while this user-space runtime provides glibc 2.28. Standard and Advanced therefore use isolated patches to expose the Transformers backends already present in the source tree.
Basic-mode performance optimization
The performance bottleneck is not the single-page operator, but "create a new Python + load all models for each task".
6.1 ONNX Runtime CPU thread configuration
| intra-op thread | Cold 2 pages | Warm 2 pages | Suggestions |
|---|---|---|---|
| 4 | 76.29 s | 10.43 s | Resource conservation |
| 8 | 31.63 s | 6.60 s | Total throughput priority |
| 16 | 33.64 s | 6.13 s | Single worker delay priority |
6.2 Torch on Z100
| Status | 2 pages take | changes |
|---|---|---|
| Cold run | 198.12 s | Model loading dominates |
| Warm run, SDPA math fallback | 2.86 s | Reuse a single loaded model |
| warm, eager formula attention | 2.64 s | About 8% faster |
mineru4-drop keeps one MinerUParser resident in each Slurm worker and disables formula SDPA on gfx906. When a 14-page document is split into 8- and 6-page chunks, the first chunk takes 148.8 seconds and the second takes 10.3 seconds.
3.4.4 and 4.0.7 controlled comparison
Both runs use the same document PDF (pages 1–2), 8 CPU cores, and 27 GB of memory. Each DCU run uses one Z100 card.
| version/path | Slurm total time taken | model stage | Judgment |
|---|---|---|---|
| 3.4.4 pipeline / CPU | 165 s | 36.95 s | Baseline |
| 4.0.7 Basic / ONNX CPU | 93 s | 37.15 s | Peripheral startup is shorter |
| 3.4.4 pipeline / Z100 | 236 s | 60.02 s | Faster Cold run DCU run |
| 4.0.7 Basic / Torch Z100 | 293 s | 71.97 s | Slower on a cold run |
- CPU model stage is almost unchanged; 4.0.7 cold-start advantage comes from import, preprocessing and output chain.
- DCU the 4.0.7 cold run is slower, but warm drops to 2.64 seconds / 2 pages.
- The versions use different output protocols, so total runtime is an application-level comparison, not a kernel-only benchmark.
MinerU 4.0.7 functional verification
| Tier | Resources / Backend | Sample | Result | Positioning |
|---|---|---|---|---|
| Flash | CPU native text | 14 pages | 69 s operation, model stage 16.8 s | Preview / Index |
| Basic | ONNX CPU / Torch Z100 | 2-page and 14-page blocks | Validated and optimized | Local high quality |
| Standard | Torch + Transformers VLM | 1 page Cold run | 160.9 s | Experimental |
| Advanced | Transformers VLM | 1-page warm run | 24.3 s | Experimental |
Standard was measured on a a cold run, while Advanced reused the VLM loaded in the same process. These results are not directly comparable in quality or speed; they show only that all four tiers can parse a document in the target environment.
MinerU Drop batch-processing interface
9.1 Adapter implementation
| remote root | version | Purpose |
|---|---|---|
mineru-drop | 3.4.4 | Stable production default |
mineru4-drop | 4.0.7 Basic/Torch | New protocol and experimental batch processing |
# 3.4.4
mineru-drop push <PDF_OR_DIRECTORY>
# 4.0.7
mineru-drop --remote-root mineru4-drop push <PDF_OR_DIRECTORY>
# Generic commands
mineru-drop [--remote-root ...] status --batch <BATCH_ID>
mineru-drop [--remote-root ...] retry --batch <BATCH_ID>
mineru-drop [--remote-root ...] fetch --batch <BATCH_ID>
Documents of up to 128 pages are processed as one batch; longer documents are split into 96-page chunks. For a small number of long documents, use one worker to avoid repeated model Cold run starts. Scale out to four workers for larger batches.
Deployment recommendations
10.1 Batch document processing
3.4.4 pipeline
has validated four-card throughput, stability, backpressure, and failure recovery.
10.2 Quick search and preview
4.0.7 Flash
requires no local large model and suits an initial scan or indexing pass.
10.3 4.x output workflow
4.0.7 Basic
CPU uses ONNX 8/16 threads; Z100 must be a resident model.
10.4 Evaluation of complex-page parsing
Standard / Advanced
Use only for manually selected small batches; exclude them from default automatic routing.
10.5 Recommended workflow
Native text; indexing only → 4.0.7 Flash
Final Markdown / JSON required → 3.4.4 pipeline
4.x schema required → 4.0.7 Basic
Complex pages; manual review → Standard / Advanced
Reproduction procedure and scope
11.1 Minimum reproduction requirements
- The login node prepares dependencies and models; the computing node runs offline.
- module must be loaded before venv.
- The main process and the spawn child process must share the glibc 2.28 launcher.
- Slurm scripts must pass real workload exit codes.
- 4.0.7 ONNX Runtime 1.20.1 uses relabeled wheel and can only run under external glibc 2.28.
11.2 Known limitations
| Limitation | Impact |
|---|---|
| No FlashAttention on gfx906 | Torch formula/VLM attention fallback; the formula model has been switched to eager |
| mineru-llama-cpp requires glibc 2.34 | the official local Standard/Advanced engine cannot be installed |
| The vendor vLLM version predates MinerU 4.0.7 | Not used as the 4.x default VLM engine |
| Cross-block paragraphs and tables | The manifest flags these cases for sampled review |
retains Triton, vLLM, glibc, failure paths, service pool implementation and all historical commands.