MinerU on Hygon Z100 DCU
This archive documents environment validation, service-pool implementation, and performance testing. It records the results, reproduction commands, measurements, and source job IDs for each approach, along with the evidence used to test the underlying assumptions.
Key findings
| Metric | Result | Evidence |
|---|---|---|
| Service-pool throughput | 955 docs/hour(4 cards, concurrency 12) | soak.log |
| Stability | 72 requests; 0 failures; 0 rejections | §10 |
| Fault recovery | Worker killed during load test; restarted automatically without losing requests | §10.3 |
| Memory leak | None (no RSS growth across 72 requests) | §10.4 |
| Backend selection | pipeline 3.9 s/page; VLM at 14.93 s/page is not recommended | §8.1 |
| FP16 / FP32 / BF16 | 16.15 / 9.01 / 5.83 TFLOPS | §3 |
Triton / torch.compile | ✕ Architecture limitation Not supported on gfx906 | §4 |
| vLLM HIP kernel | ✓ Available PagedAttention verified in testing | §5 |
| CPU queue | Equivalent output; 8.9× slower at steady state | §8.3 |
| Reproduction after environment drift | ✓ Passed 14-page pipeline run; exit code 0 | §14 |
MinerU can run in production on Z100 using the pipeline backend. Triton is unavailable on this architecture, but MinerU does not require it.
Background and validation scope
2.1 Initial environment status
The prior work set up the environment and produced a V2 report claiming that “all tests passed.” A review found one critical gap:
ssh -p <SSH_PORT> <CLUSTER_HOST> 'ls ~/.cache/modelscope/hub ~/models ~/output 2>/dev/null'
# all directories are empty
MinerU had never parsed a document. The V2 report's “end-to-end test” measured only magika file-type detection—the first preprocessing step in MinerU, not document parsing. No models had been downloaded.
2.2 Reusable software and data
| Item | Path / version |
|---|---|
| venv | <USER_HOME>/mineru-venv-py310(Python 3.10.14) |
| torch | 2.7.1+das.opt1.dtk2604 |
| MinerU | Version 3.4.4 (editable install; source patched for magika compatibility) |
| glibc 2.28 | <USER_HOME>/glibc-2.28-rpm/ (path used at the time; relocated during revalidation, see §14) |
| Wrappers | dcu-python / mineru / mineru-api |
2.3 Cluster resources and runtime constraints
cat ~/<SKILL_CONFIG>
| Constraint | Value |
|---|---|
| Memory limit | cpus-per-task × 3569 MB |
| DCU queue | kshdnormal, QOS enforces --gres=dcu:1 |
| CPU queue | kshcnormal, several nodes; no GRES restriction |
| External network on compute nodes | ❌ None |
/tmp | Not shared across nodes |
Hardware capabilities and software environment
3.1 Initial probe results and limitations
Probe job <JOB_ID>:
sbatch ~/probe_caps.slurm # see cluster/probe_caps.py
torch 2.7.1 devices 4
arch: gfx906:sramecc-:xnack- mem GB: 17.2
[OK] bf16 matmul: supported=True maxerr=0.250
[OK] fp16 matmul: 4.85 TFLOPS ← suspicious: slower than FP32
[OK] fp32 matmul: 8.91 TFLOPS
[FAIL] triton kernel: FileNotFoundError: '.../gcc-11.2.0-install/bin/gcc'
[OK] SDPA: flash_avail=True mem_eff=True
[FAIL] torch.compile: FileNotFoundError: (same error)
[OK] torch._int_mm: ok (512, 512)
[OK] onnxruntime: 1.16.3 providers=['AzureExecutionProvider','CPUExecutionProvider']
FP16 performance below FP32 was unexpected—the investigation found that the benchmark had no warm-up.
Benchmark method after warm-up was added
Job <JOB_ID>, with five warm-up iterations followed by 30 timed iterations:
def bench(dt, n=2048, iters=30):
a = torch.randn(n, n, device="cuda", dtype=dt)
b = torch.randn(n, n, device="cuda", dtype=dt)
for _ in range(5): c = a @ b # ← warm-up
torch.cuda.synchronize(); s = time.time()
for _ in range(iters): c = a @ b
torch.cuda.synchronize()
return 2 * n**3 / ((time.time() - s) / iters) / 1e12
fp32: 9.01 TFLOPS
fp16: 16.15 TFLOPS ← 1.8× FP32
bf16: 5.73 TFLOPS
--- force the rocBLAS/hipBLAS backend ---
fp16: 16.12 TFLOPS ← unchanged; rocBLAS was already in use
bf16: 5.83 TFLOPS
3.3 Final capability matrix
| Capability | Result | Notes |
|---|---|---|
| FP16 GEMM | 16.15 TFLOPS | Use FP16 for inference |
| FP32 GEMM | 9.01 TFLOPS | |
| BF16 GEMM | 5.83 TFLOPS | Software emulation; slowest option |
torch._int_mm | ✓ | INT8 available |
| SDPA | ⚠ Degraded | Falls back to the math backend; no Flash or memory-efficient attention |
| Multi-GPU | ✓ 4 cards | 16 GB per card (the report states 17.2 GB) |
| hipBLASLt | ✕ No gfx906 kernel | The library contains kernels only for gfx928; runtime warnings are emitted |
| FP8 / MFMA | ✕ | Vega20 architecture limitation |
Benchmarks involving one-time costs (kernel tuning, JIT compilation, or model loading) must include warm-up; otherwise, the conclusion may be completely reversed.
Triton build failure analysis
This case shows why environment failures must be traced across layers: the initial error did not identify the underlying cause.
4.1 Incorrect GCC path
FileNotFoundError: '<USER_HOME>/software/gcc-11.2.0-install/bin/gcc'
This directory appeared in the user's PATH, but it did not exist:
echo $PATH | tr ':' '\n' | grep gcc
# <USER_HOME>/software/gcc-11.2.0-install/bin ← invalid path
ls <USER_HOME>/software/gcc-11.2.0-install/bin/gcc
# ls: cannot access ...: No such file or directory
Triton initially appeared unavailable; investigation showed that the configured GCC path pointed to a nonexistent directory.
4.2 System GCC version and library-path issue
After correcting the path to the existing gcc-11.2.0:
cc1: error while loading shared libraries: libisl.so.15: cannot open shared object file
Compilers installed on the cluster, tested individually:
echo "int main(){return 0;}" > /tmp/t.c
for cc in <SOFTWARE_ROOT>/compiler/gcc-11.2.0/bin/gcc \
<SOFTWARE_ROOT>/compiler/gcc-12.2.0/bin/gcc \
<SOFTWARE_ROOT>/compiler/gcc-13.3.0/bin/gcc \
/opt/rh/devtoolset-7/root/usr/bin/gcc \
<SOFTWARE_ROOT>/compiler/rocm/dtk-26.04/llvm/bin/clang; do
$cc /tmp/t.c -o /tmp/t.out 2>/dev/null && echo "OK $cc" || echo "FAIL $cc"
done
| Compiler | Result |
|---|---|
| gcc-11.2.0 | ✕ Missing libisl.so.15 |
| gcc-12.2.0 | ✓ |
| gcc-13.3.0 | ✕ Missing libisl.so.15 |
| devtoolset-7 | ✓ |
| DTK clang | ✓ |
4.3 Root cause: unsupported gfx906 target
After switching to GCC 12.2.0 (job <JOB_ID>):
loc("probe_triton.py":4:0): error: unsupported target: 'gfx906'
RuntimeError: PassManager::run failed
4.4 Validation against Hygon's Triton build
Hygon also publishes a DCU-specific Triton build, which warranted a separate test:
curl -sL "https://download.sourcefind.cn:65024/directlink/4/triton/DAS1.8/"
# triton-3.1.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl
Install it in an isolated directory to avoid affecting the working environment:
cp triton-3.1.0+...manylinux_2_28_x86_64.whl triton-3.1.0+...manylinux2014_x86_64.whl
pip install --no-deps --target=<USER_HOME>/dcu-triton-test
The first run could not find libgcvm.so.17git. The library was in dtk-26.04/dcc/gcvm/lib, which was missing from the wrapper's library path. Adding that directory to dcu-python's ALL_LIBS resolved the issue (job <JOB_ID>):
triton: 3.1.0 from <USER_HOME>/dcu-triton-test/triton/__init__.py
arch: gfx906:sramecc-:xnack-
[FAIL] RuntimeError: PassManager::run failed
error: unsupported target: 'gfx906' ← same result as the PyPI build
Source inspection confirmed the architecture allowlist:
grep -ohE "gfx9[0-9]{2}[a-z]*" triton/backends/amd/compiler.py | sort -u
# gfx928 gfx936 gfx940 gfx941 gfx942
grep -ohE "gfx9[0-9]{2}[a-z]*" triton/backends/hcu/compiler.py | sort -u
# gfx928 gfx936 gfx938 gfx940 gfx941 gfx942
gfx906 is not listed.
strings libtriton.so | grep gfx906 A search does find gfx906. That entry is in LLVM upstream's AMDGPU target list. LLVM recognizing an architecture does not mean Triton supports it. Test the installed build; do not infer support from strings in the binary.
Neither Triton 3.x (the PyPI build nor Hygon's DCU build) supports gfx906,torch.compile so Triton-dependent features are unavailable This is an architectural limitation; configuration changes cannot resolve it.
vLLM porting feasibility
5.1 Backend capability boundaries
First, distinguish two easily confused terms:
- VLM = Vision-Language Model, a type ofmodel(MinerU2.5-Pro-2605-1.2B is one example)
- vLLM is an inference engine that provides PagedAttention.
MinerU has two layers of options, differing by one letter:
5.2 gfx906 kernel architecture validation
Hygon also provides a DCU build of vLLM:
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vllm/DAS1.8/"
# vllm-0.11.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl
Inspect the packaged code objects:
strings vllm/_C.abi3.so | grep -oE "gfx[0-9]{3,4}" | sort -u
# gfx906 gfx926 gfx928 gfx936 gfx938
strings vllm/_moe_C.abi3.so | grep -oE "gfx[0-9]{3,4}" | sort -u
# gfx906 gfx926 gfx928 gfx936 gfx938
Hygon excludes gfx906 from Triton's supported targets but retains it in vLLM's HIP kernels.
5.3 Runtime test (job <JOB_ID>)
from vllm import _custom_ops as ops
# PagedAttention is vLLM's core attention operation
ks = torch.tensor(1.0, device="cuda"); vs = torch.tensor(1.0, device="cuda")
ops.paged_attention_v1(out, q, kc, vc, nkvh, 1.0/(hd**0.5),
btab, slen, bs, bs*2, None, "auto", ks, vs, 0,0,0,64,0)
arch: gfx906:sramecc-:xnack-
[OK] import vllm 0.11.0
[OK] vllm._C loaded
[OK] rms_norm (HIP): maxerr=0.0010
[OK] rotary_embedding (HIP): ran
[OK] paged_attention_v1 (HIP) ★: out(4, 8, 64) finite=True
5.4 Conclusion and remaining work
vLLM has a working execution path on Z100—all core HIP kernels work, the opposite of the Triton result.
Unresolved: platform detection fails (is_rocm: False). vLLM uses import amdsmi to detect ROCm:
# vllm/platforms/__init__.py:109
import amdsmi
amdsmi.amdsmi_init()
but DTK 26.04 provides only the older rocm_smi interface (.hyhal/rocm_smi/bin/rsmiBindings.py). Completing inference requires an amdsmi shim or a patch to rocm_platform_plugin().
No benefit for MinerU: vLLM's strengths are PagedAttention and continuous batching for high-concurrency, long-form generation. MinerU's VLM workload uses a single-page image and short output, so it does not benefit from these features. The vLLM path is still worth pursuing for an LLM inference service on Z100.
Parsing workflow failures and fixes
Each new error showed that diagnosis had progressed to the next layer. The issues below are listed in the order encountered:
6.1 Model not available
modelscope imports torch and fails on the login node, which has no DCU driver. Use huggingface_hub (pure Python) with hf-mirror instead:
# Check endpoint reachability
curl -s -o /dev/null -w "%{http_code}" https://hf-mirror.com/ # 200
curl -s -o /dev/null -w "%{http_code}" https://huggingface.co/ # blocked from this network
# cluster/dl_models.py
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
os.environ["HF_HOME"] = "<USER_HOME>/mineru-models/hf"
from huggingface_hub import snapshot_download
for repo in ["opendatalab/MinerU2.5-Pro-2605-1.2B",
"opendatalab/PDF-Extract-Kit-1.0"]:
snapshot_download(repo, max_workers=8)
This downloads 20 GB. Set the local path:
{
"models-dir": {
"pipeline": "<USER_HOME>/mineru-models/models/OpenDataLab--PDF-Extract-Kit-1.0/snapshots/master",
"vlm": "<USER_HOME>/mineru-models/models/OpenDataLab--MinerU2.5-Pro-2605-1.2B/snapshots/master"
},
"model-source": "local",
"device-mode": "cuda"
}
6.2 Resolve libssl.so.3 not found
File "mineru/cli/fast_api.py", line 19, in
import uvicorn
ImportError: libssl.so.3: cannot open shared object file
Error: Local mineru-api exited before becoming healthy.
Root cause: The MinerU CLI launches a mineru-api child process with the system sys.executable, bypassing the dcu-python wrapper.
Workaround used at the time: Call the do_parse() Python API instead of the CLI. Permanent service fix: Add /usr/local/lib64 to the wrapper's ALL_LIBS.
6.3 operator torchvision::nms does not exist
torchvision 0.21.0 from PyPI is incompatible with the DCU Torch 2.7.1 ABI.
The DCU torchvision build is published in the sourcefind vision/ directory, not torchvision/:
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vision/DAS1.8/"
# torchvision-0.22.0+das.opt1.dtk2604.torch271-cp310-...whl ← built for torch 2.7.1
Because the system glibc is 2.17, pip rejects the manylinux_2_28 wheel. Retagging lets pip install it; the program itself already runs under the user-space glibc 2.28 loader:
W=torchvision-0.22.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl
cp $W ${W/manylinux_2_28/manylinux2014}
pip install --no-deps --force-reinstall ${W/manylinux_2_28/manylinux2014}
6.4 Missing OCR/layout dependencies
pip install shapely pyclipper omegaconf einops ftfy
6.5 Resolve BrokenProcessPool
concurrent.futures.process.BrokenProcessPool:
A process in the process pool was terminated abruptly
Root cause: MinerU renders PDFs using a spawn process pool:
# mineru/utils/pdf_image_tools.py:160
if start_method != "spawn":
return ProcessPoolExecutor(max_workers=max_workers,
mp_context=multiprocessing.get_context("spawn"))
The spawn method starts a fresh interpreter through sys.executable, which loses the configured library paths.
Fix——sitecustomize.py, imported automatically at Python startup so the change applies globally:
# /lib/python3.10/site-packages/sitecustomize.py
import os
_w = "<USER_HOME>/mineru-venv-py310/bin/dcu-python"
if os.path.exists(_w):
import multiprocessing
multiprocessing.set_executable(_w)
try:
multiprocessing.get_context("spawn").set_executable(_w)
except Exception:
pass
Verification:
dcu-python -c "
import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor
def f(x): return x*x
if __name__=='__main__':
with ProcessPoolExecutor(2, mp_context=mp.get_context('spawn')) as ex:
print(list(ex.map(f,[1,2,3])))"
# [1, 4, 9]
Any script that calls MinerU must include an if __name__ == "__main__": guard. A spawned worker imports the main module; without the guard, the script runs again, producing duplicate [info] output and then a BrokenProcessPool error. The MinerU documentation does not mention this requirement.
6.6 Resolve Unsupported model IR version: 10
onnxruntime.capi.onnxruntime_pybind11_state.Fail:
Load model from .../PP-LCNet_x1_0_table_cls.onnx failed:
Unsupported model IR version: 10, max supported IR version: 9
onnxruntime 1.16.3 supports IR up to v9, but MinerU's table-classification model uses IR v10.
The V2 report concluded that ONNX Runtime ≥1.17 was incompatible because it was published only for manylinux_2_28. The wheel can in fact be retagged using the method in §6.3; this step is required.
curl -sL "https://pypi.tuna.tsinghua.edu.cn/simple/onnxruntime/" \
| grep -oE 'href="[^"]*onnxruntime-1\.20\.1-cp310-cp310-manylinux_2_27_x86_64[^"]*"'
curl -sL -o ort.whl "https://pypi.tuna.tsinghua.edu.cn/packages/63/47/.../onnxruntime-1.20.1-...whl"
cp ort.whl onnxruntime-1.20.1-cp310-cp310-manylinux2014_x86_64.whl
pip install --no-deps --force-reinstall onnxruntime-1.20.1-cp310-cp310-manylinux2014_x86_64.whl
6.7 Additional fixes
| Issue | Fix |
|---|---|
onnxruntime emits pthread_setaffinity an error | Set OMP/ORT/OPENBLAS/MKL_NUM_THREADS=4 |
libgcvm.so.17git is missing | Add dtk-26.04/dcc/gcvm/lib to the ALL_LIBS |
No module named mineru.cli.api | In version 3.4.4 the module is named fast_api; update the mineru-api wrapper |
Output quality validation
Job <JOB_ID>; 23-page Mooncake paper, 0.58 MB.
dcu-python parse_api.py ~/sourcecode/Mooncake/Mooncake-v3.pdf ~/parse-out/api-pipeline pipeline
[RESULT] backend=pipeline elapsed=146.8s
[OUT] .../Mooncake-v3.md chars=81072
processing-window multi-file infer finished, cost: 90.52, speed: 0.254 page/s
Output files:
Mooncake-v3.md 81 KB Markdown source
Mooncake-v3_content_list.json 121 KB structured content
Mooncake-v3_middle.json 2.1 MB intermediate representation
Mooncake-v3_layout.pdf 835 KB layout visualization
Mooncake-v3_span.pdf 832 KB span visualization
images/ 26 files
Checks performed (not just relying on a zero job exit code):
| Item | Result | Verification |
|---|---|---|
| Markdown | 81,072 characters | Heading hierarchy, superscript author symbols ♠♡, and abstract structure all correct |
| Tables | 3 HTML tables | Sampled LRUCache/LFUCache rows; values and rows are complete |
| Images | 23 references / 26 files | images/ Files present on disk |
| Headings | 31 | Hierarchy matches the source |
| Inline equations | 17 LaTeX expressions | $T_{queue}$、$\mathrm{MLP}$、$T_{prefill}$ |
| Display equations | None | The source uses inline math; no missed display equations |
Excerpt from the first page:
# Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Ruoyu Qin♠♡1 Zheming Li♠1 ... ♠Moonshot AI ♡Tsinghua University
## Abstract
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI...
Searching only for $$ display equations returned no matches and nearly led to the wrong conclusion that formula recognition had failed. The paper uses inline math: middle.json contains 38 inline_equation entries, and MFR identifies 19 regions. Choose checks that match the content; the absence of one format does not prove a feature is missing.
Performance tests and results
8.1 Backend comparison
| Backend | Job | Throughput | Assessment |
|---|---|---|---|
| pipeline (DCU) | <JOB_ID> | 3.9 s/page(90.5 s / 23 pages) | Selected for production |
| pipeline (CPU) | <JOB_ID> | 10.1 s/page (231.3 s) | Equivalent output; see §8.3 |
| VLM (transformers) | <JOB_ID> | 14.93 s/page | Timed out after 50 minutes; incomplete |
VLM is 3.8× slower and has no acceleration path on gfx906. Triton is unavailable, so the vllm-engine path cannot run. VLM is not recommended for production.
Confirming that VLM was computing rather than stalled:
ssh <COMPUTE_NODE> "cat /proc/<PID>/status | grep -E '^State|^Threads'; ls /proc/<PID>/fd | grep -c kfd"
# State: R (running) Threads: 44 kfd handles: 1
8.2 Model loading dominates runtime
Share of runtime spent loading the model:
Model loading accounts for about 85% of the runtime. This rules out the “one job per document” design: 85% of the allocated time would be spent reloading the same models. Keep workers resident and process documents in batches.
8.3 CPU-queue output equivalence
Byte-for-byte comparison of the output:
diff ~/parse-out/cpu/Mooncake-v3/auto/Mooncake-v3.md \
~/parse-out/api-pipeline/Mooncake-v3/auto/Mooncake-v3.md
# Five-line diff:
# 98,99c98
# < ## Listing 1: Request samples.
# ---
# > Listing 1: Request samples.
81,074 vs 81,072 characters; the only difference is one heading level. The CPU queue is a viable production path, not merely a fallback.
However, a single-document comparison substantially understates the performance gap:
| Single document | Batch steady state | |
|---|---|---|
| DCU | 146.8 s | 19.6 s |
| CPU | 315.4 s | 173.8 s |
| Ratio | 2.1× | 8.9× |
For one document, model loading dominates both paths, making CPU and DCU runtimes appear similar. This understates the steady-state throughput gap by about four times.
On the CPU path, formula recognition is the main cost: MFR runs at 5.86 it/s on DCU versus 3.85 s/it on CPU (about 22× slower).
8.4 Batch throughput (without HTTP)
DCU job <JOB_ID>; CPU job <JOB_ID>:
dcu-python throughput.py ~/pdf-corpus ~/parse-out/tput pipeline 4
| Configuration | Including cold start | Steady state | Steady-state time per document |
|---|---|---|---|
| 4 DCU | 233 docs/h | 735 docs/h | 19.6 s |
| 4 CPU worker | 63 docs/h | 83 docs/h | 173.8 s |
With load balanced across four workers, runtimes were 176.4 / 177.4 / 185.3 / 185.7 s, a spread of less than 6%.
Bug found: workers did not exit after batch processing and remained stuck, each of the four processes retaining 3.8 GB RSS. MinerU's _get_pdf_render_executor() creates a persistentpool; its non-daemon child processes prevent the parent from exiting. Fix:
from mineru.utils.pdf_image_tools import shutdown_pdf_render_executor
shutdown_pdf_render_executor() # called during worker shutdown
After the fix, the CPU batch job exits normally and prints BENCH_DONE and RC=0 exits cleanly.
8.5 Queue configuration comparison
Cluster load during the measurements:
sinfo -p kshdnormal -o "%t %D" -h | sort # alloc 735 / idle 2 (fully loaded)
sinfo -p kshcnormal -o "%t %D" -h | sort # idle 1106 (idle)
Total throughput = throughput per node × number of available nodes:
- A few idle DCU nodes × 735 ≈ 1,470 docs/h
- Several idle CPU nodes × 83 ≈ 91,663 docs/h
The order-of-magnitude difference comes from node availability, not compute capacity. Use the CPU queue for large offline batches; reserve DCUs for latency-sensitive online requests and formula-heavy documents.
Service-pool architecture and implementation
9.1 Network reachability and service architecture
Job <JOB_ID>. First verify whether compute nodes can serve requests:
# Start the service on a compute node and write its endpoint
IP=$(hostname -I | awk '{print $1}'); PORT=$((30000 + RANDOM % 5000))
dcu-python svc_hello.py $PORT &
echo "$IP:$PORT" > ~/svc_endpoint.txt
# Access it from the login node
curl --max-time 8 "http://$(cat ~/svc_endpoint.txt)/"
# MINERU_SVC_ALIVE on <COMPUTE_NODE> ← reachable
| Test | Result |
|---|---|
Compute node binds to 0.0.0.0 Self-test | ✓ |
| Login node → compute-node port | ✓ |
| Compute node → external network | ✕ (expected) |
Cluster-internal network connectivityis available; this is separate from the lack of external access on compute nodes. The service pool can therefore expose a reachable service endpoint.
9.2 Service-pool architecture
Design rationale(Each decision is based on measurements, not assumptions):
| Decision | Rationale |
|---|---|
| One process per card | MinerU models are not thread-safe; concurrent requests in one process serialize |
HIP_VISIBLE_DEVICES Pin each worker to a card | Isolates GPU memory; an OOM in one worker does not affect the others |
| Limit each worker to one concurrent request | Established by load testing; see §10.1 |
| Use file modification time as a heartbeat | Workers register and deregister themselves; the gateway scans the directory, with no separate service-discovery component |
| Retry failed requests on another worker | Parsing is idempotent, so retrying is safe |
9.3 MinerU API endpoints
grep -n "@app\.\(get\|post\)" -A3 mineru/cli/fast_api.py
| Route | Purpose |
|---|---|
POST /file_parse | Synchronous parsing |
POST /tasks | Asynchronous submission |
GET /tasks/{id} /tasks/{id}/result | Asynchronous status query |
GET /health | Health check (for gateway) |
9.4 Key implementation excerpts
Gateway scheduling (scripts/mineru_gateway.py):
def _acquire(exclude=()):
"""Reserve a slot on the least-loaded healthy worker, waiting if all busy."""
deadline = time.time() + SLOT_WAIT_TIMEOUT
with _cv:
while True:
cands = [(v["inflight"], v["served"], u)
for u, v in _workers.items()
if v["healthy"] and u not in exclude
and v["inflight"] < MAX_INFLIGHT] # rate limit
if cands:
cands.sort()
url = cands[0][2]
_workers[url]["inflight"] += 1
return url
if time.time() >= deadline: return None
_cv.wait(2.0) # queue instead of rejecting
Retry after failure:
for attempt in range(MAX_ATTEMPTS):
url = _acquire(exclude=tuple(tried)) # exclude workers that have already failed
...
except urllib.error.HTTPError as e:
if e.code < 500: # client error (4xx): do not retry
return
_release(url, False)
tried.append(url)
Worker supervision (cluster/mineru_worker.sh):
while true; do
mineru-api --host 0.0.0.0 --port $PORT &
APID=$!
# Register only after health checks pass; do not route traffic during model loading
for i in $(seq 1 90); do
curl -sf --max-time 3 "http://127.0.0.1:$PORT/health" >/dev/null && READY=1 && break
sleep 2
done
[ "$READY" = 1 ] && echo "{\"url\":\"$URL\",...}" > "$REGFILE"
# Heartbeat: touch while healthy; the gateway checks the modification time
while kill -0 $APID <PID>>/dev/null; do
curl -sf --max-time 5 ".../health" >/dev/null && touch "$REGFILE"
sleep 10
done
rm -f "$REGFILE"; sleep 5 # restart after a crash
done
9.5 Startup
sbatch ~/mineru_pool.slurm # cluster/mineru_pool.slurm
[t+95s] healthy workers: 4
=== POOL READY ===
{"pool_size": 4, "healthy": 4, "max_inflight_per_worker": 1, ...}
Verified from the login node:
curl -s "http://<SERVICE_ENDPOINT>/pool/status" | python3 -m json.tool
Service-pool performance and stability tests
10.1 Performance comparison between gateway versions
Same workload: 24 requests at concurrency 8
python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 8 2
| Metric | v1 (no concurrency limit or retry) | v2 (production version) |
|---|---|---|
| Error rate | 8.3% (2 × 502) | 0.0% |
| Throughput | 374 docs/h | 440 docs/h (+18%) |
| Load distribution | 11 / 4 / 5 / 2 | 6 / 5 / 7 / 6 |
| p50 | 27.2 s | 38.0 s |
| p90 | 111.9 s | 121.7 s |
Raw output:data/loadtest.log、data/loadtest2.log
Concurrency limits improved throughput—MinerU models are not thread-safe, so requests within a worker are processed serially. Sending more work to a busy worker only creates an internal queue. Load skew fell from 5.5:1 to 1.4:1.
The increase in p50 latency is expected: v1's lower p50 was an artifact of a few requests reaching idle workers; p90 latency was higher and 8.3% of requests failed. v2 queues requests fairly across the pool.
10.2 Overload test at three times worker capacity
48 requests at concurrency 12:
python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 12 4
total=48 ok=48 failed=0 error_rate=0.0%
wall=180.9s throughput=955 docs/hour concurrency=12
latency p50=32.4s p90=70.4s p99=131.1s min=11.9s max=131.1s
per-worker distribution: 11 / 11 / 14 / 12
Pool status during the load test:
"in_flight": 4 "waiting_for_slot": 8 "queue_peak": 8 "retries": 0 "rejected": 0
Backpressure worked as designed: four requests were allowed in flight; the other eight queued instead of failing.The pool slowed under overload without dropping requests.
Throughput rose from 440 to 955 docs/hour as concurrency increased from 8 to 12, showing that the pipeline was not saturated at concurrency 8. The 955 docs/hour resultexceeded the offline batch rate of 735because static per-card sharding leaves workers idle, while the service pool dispatches work dynamically to whichever slot becomes available.
10.3 Worker recovery under load
PID=$(python3 -c "import json;print(json.load(open('.../<COMPUTE_NODE>-dcu3.json'))['pid'])")
ssh <COMPUTE_NODE> "kill -9 $PID" # PID <PID>
After 50 seconds:
before: "healthy": 3 "total_served": 5 "total_failed": 0
after: "healthy": 4 "total_served": 16 "total_failed": 0
Confirmed worker replacement by PID change and logs:
cat ~/mineru-registry/<COMPUTE_NODE>-dcu3.json
# {"url":"http://<SERVICE_ENDPOINT>",...,"pid":<PID>} ← 原 <PID>
grep -a dcu3 pool-*.err | grep -E "restart|READY|exited"
# [worker dcu3] exited, restarting in 5s (total restarts: 1)
# [worker dcu3] starting on http://<SERVICE_ENDPOINT> (restart #1)
# [worker dcu3] READY -> registered http://<SERVICE_ENDPOINT>
Recovery was tested under real load;total_failed there were zero failures and no lost requests.
10.4 Memory stability across 72 requests
ssh <COMPUTE_NODE> "ps -u <USER> -o rss,cmd --sort=-rss | grep -a fast_api | awk '{print \$1/1024}'"
| w0 | w1 | w2 | w3 | |
|---|---|---|---|---|
| Before | 4839 | 4780 | 4754 | 4744 |
| After | 4837 | 4773 | 4762 | 4742 |
No growth. The process count remained at four, with no buildup of spawned children.
10.5 Summary
Approaches not adopted and operating limits
Failed paths and their costs are documented to avoid repeating the same work.
11.1 Embedding library paths in the ELF with patchelf ✕ Failed
Rationale: The ld-linux wrapper covers only the top-level process. Embedding library paths in the Python binary with patchelf might let every child process inherit them, which would be cleaner in theory.
patchelf --set-interpreter $GLIBC_DIR/ld-linux-x86-64.so.2 python3.10
Inconsistency detected by ld.so: dl-call-libc-early-init.c: 37:
_dl_call_libc_early_init: Assertion `sym != NULL' failed!
Even with an added rpath, the process still segfaulted; ldd also crashed:
patchelf --set-interpreter ... --force-rpath --set-rpath "$G:/usr/local/lib64:/usr/lib64" python3.10
./python3.10 -c "print(1)" # Segmentation fault
ldd python3.10 # Segmentation fault
Isolation tests:
| Test | Result |
|---|---|
| Change only the rpath; keep the system interpreter | ✓ Passed |
| Change only the interpreter | ✕ Assertion failed |
| Change both | ✕ segfault |
The glibc 2.28 ld.so loader cannot be combined with the system glibc 2.17 libraries, so the change was reverted from backup. Do not patch the ELF interpreter across this glibc version boundary; use a wrapper and sitecustomize.py.
11.2 VLM backend ⚠ Runs, but is not practical
At 14.93 s/page, the 23-page run did not finish within the 50-minute limit. No acceleration path is available on gfx906: Triton is unavailable, so the vllm-engine / lmdeploy engine paths cannot run; only single-sequence inference through Transformers is available. The pipeline backend already produces higher-quality output and is 3.8× faster.
11.3 Full vLLM inference ⏸ Incomplete, not failed
Core HIP kernels work in testing (§5), but platform detection is blocked at import amdsmi. DTK 26.04 provides only the older rocm_smi. A shim or patch is needed.The path is viable; the integration work is unfinished.
11.4 Failed approaches recorded in the V2 report
| Approach | Result | Reason |
|---|---|---|
| Install onnxruntime with Conda | ✕ | Channel access timed out |
| Singularity container | ✕ | Docker Hub was blocked |
| Python 3.8 | ✕ | MinerU requires Python ≥3.10 |
| Force-install Python 3.12 + onnx 1.17 | ✕ (at the time) | This work resolved the same class of issue by retagging the wheel |
Assumptions tested and revised
Three course corrections during the work changed the direction of the conclusions. They are recorded here because the method matters as much as the results.
Without warm-up, FP16 measured 4.85 TFLOPS, below FP32 at 8.91 TFLOPS, suggesting that FP16 offered no benefit on gfx906. After warm-up, FP16 reached 16.15 TFLOPS, 1.8× the FP32 result.
Warm up any benchmark that includes one-time costs.
The initial assessment tied vLLM availability on gfx906 to Triton support, without distinguishing their dependency layers. vLLM's PagedAttention uses hand-written HIP kernels and does not depend on Triton. Inspection confirmed that Hygon's _C.abi3.so contains kernels compiled for gfx906, and the paged_attention_v1 operation ran successfully.
Do not infer one component's support from another component's failure, even when they appear related. Validate each layer separately.
For one document, 315.4 s vs 146.8 s suggested a 2.1× difference, leading to the claim that “the CPU queue is only slightly slower.” Batch steady-state measurements were 173.8 s vs 19.6 s, an 8.9× difference.
A single-document run is dominated by model loading on both systems. Since model-loading times are similar on CPU and DCU, they mask the true throughput gap by a factor of four.
Compare steady-state results when startup costs are fixed overhead.
The initial check searched only for display equations delimited by $$ and found none. The paper instead uses inline math throughout: 17 $...$ expressions appear in the rendered text, and middle.json contains 38 inline_equation regions.
Choose checks that match the content; the absence of one format does not prove a feature is missing.
During environment revalidation on 2026-09-27, the top-level Python process used the user-space glibc 2.28 loader, but the local API process launched by the MinerU CLI used the system glibc 2.17 and reported GLIBC_2.27 not found. The wheel was valid; the runtime environment had not been propagated across the process boundary.
After loading the GCC 12.2 and DTK 26.04 modules, starting the main process with the same loader, and making multiprocessing children inherit that entry point, the 14-page pipeline completed and returned 0.
Check library compatibility across the full process tree. A successful import in the parent does not prove that the CLI, service process, and spawned workers use the same runtime.
Deployment and runtime configuration
The deployment is bare metal, outside Slurm.
13.1 Runtime compatibility and operating-system requirements
The cluster's ld-linux wrapper was created to work around CentOS 7's glibc 2.17, but it leaves child-process blind spots (§6.2 and §6.5).With bare-metal deployment, the operating system is under your control, so this issue can be avoided.
The DCU Torch build requires GLIBC_2.25:
objdump -T torch.libs/libevent_core*.so* | grep -oE "GLIBC_[0-9.]+" | sort -u -V | tail -3
# GLIBC_2.14
# GLIBC_2.17
# GLIBC_2.25 ← actual requirement
Use sort -V for version-aware sorting. A lexicographic sort can place GLIBC_2.25 before GLIBC_2.9, leading to the false conclusion that 2.9 is the newest required version.
The Torch wheels tagged manylinux_2_17 and manylinux_2_28 have the same MD5 hash, 7b098c2f9c036fcdb9ae244228bac822. Only their tags differ; retagging does not remove the actual glibc requirement. Use Rocky Linux 8+, Ubuntu 20.04+, or CentOS 8+ to avoid the need for glibc wrappers and wheel retagging.
13.2 Deployment
./mineru-baremetal-deploy.sh preflight # Check glibc / DTK / virtualenv / models
sudo ./mineru-baremetal-deploy.sh install # Generate the systemd unit
sudo ./mineru-baremetal-deploy.sh start
./mineru-baremetal-deploy.sh status
systemd handles startup and restarts the service after a crash with Restart=always. Set TimeoutStartSec=300 to allow time for model loading. The registration script registers a worker only after its health checks pass, so traffic is not sent to a worker that is still loading the model.
13.3 Requests
curl -X POST http://:8000/file_parse \
-F "files=@paper.pdf" -F "backend=pipeline"
curl http://:8000/pool/status
The X-MinerU-Worker response header identifies the worker that handled the request; X-MinerU-Attempt records the retry attempt.
13.4 Monitoring
| Metric | Alert condition |
|---|---|
healthy | < Target threshold: alert |
waiting_for_slot | Sustained > more than 2× the worker count → scale out |
total_failed | Alert on abnormal growth |
retries | A sudden increase indicates an unstable worker |
rejected | Any value above 0 requires immediate investigation |
13.5 Scaling out
Scale horizontally without changing the gateway: deploy workers on new machines (without a gateway), point the registry to the same shared-storage path (NFS/GPFS), and the gateway will discover them automatically.
The maximum workers on one node depends on its card count. If CPU work (such as layout analysis and OCR preprocessing) becomes the bottleneck, reduce *_NUM_THREADS per-worker concurrency or the worker count.
Environment recovery and reproduction
On 2026-09-27, after the original experiment directory had been archived and relocated, the runtime closure was audited from scratch and an end-to-end pipeline parse was completed on a 14-page PDF.
The Kunshan module environmentdoes not provide a separate glibc module. Reproduction requires the GCC 12.2 and DTK 26.04 modules, a user-space glibc 2.28 loader, the correct local model configuration, and the same Python entry point for the parent and multiprocessing children.
14.1 Reproduction scope and evidence
| Layer | Validation | Evidence |
|---|---|---|
| Login node | No standalone glibc module; GCC 12.2 and DTK 26.04 are available | module avail/show |
| Compute node | PyTorch import, device enumeration, and FP32/FP16/BF16 GEMM | Controlled single-card probe |
| Process tree | Main Python process and PDF-render worker use the same loader | Multiprocessing test |
| End to end | 14-page PDF → Markdown / JSON / images / rendered PDF | Pipeline exit code 0 |
The public edition omits usernames, cluster addresses, node names, actual job IDs, and real paths. It retains versions, error types, command structure, performance figures, and acceptance criteria.
14.2 Environment audit
Do not rerun old commands blindly. First check whether the wrapper, configuration, and model paths still exist:
module -t avail 2>&1 | grep -Ei '(^|/)glibc([/-]|$)|compat-glibc'
# Expected: no output; glibc 2.28 is not provided by a module.
for f in \
<VENV>/bin/dcu-python \
<VENV>/bin/mineru \
<VENV>/bin/mineru-api \
<VENV>/lib/python3.10/site-packages/sitecustomize.py
do
echo "=== $f ==="
sed -n '1,80p' "$f"
done
test -x <GLIBC_ROOT>/lib64/ld-linux-x86-64.so.2
test -x <VENV>/bin/python3.10
test -f <MODEL_CONFIG>
test -d <PIPELINE_MODEL_ROOT>
This revalidation found that the four old entry points still referenced an archived glibc directory. The model weights were still present, but the configuration file had moved to a new scripts directory.The environment may still be intact even when its entry-point paths are stale.
14.3 Module environment
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VENV>/bin/activate
module list
# compiler/gcc/12.2.0
# compiler/dtk/26.04
gcc --version | head -1
# gcc (GCC) 12.2.0
echo "$ROCM_PATH"
# <SOFTWARE_ROOT>/compiler/rocm/dtk-26.04
The DTK module supplies HIP and math libraries, while the GCC module supplies a newer libstdc++.so.6. Neither replaces glibc. Loading only DTK, or adding only glibc to --library-path, can still result in CXXABI_1.3.8 not found.
14.4 A single Python entry point
Avoid maintaining several slightly different wrappers. Create one entry point and use it for the CLI, service, and spawned workers:
#!/bin/bash
set -u
GLIBC_ROOT="${GLIBC_ROOT:-<USER_HOME>/softwares/runtime/glibc-2.28}"
GCC_ROOT="${GCC_ROOT:-<SOFTWARE_ROOT>/compiler/gcc-12.2.0}"
DTK_ROOT="${DTK_ROOT:-<SOFTWARE_ROOT>/compiler/rocm/dtk-26.04}"
VENV="${VENV:-<USER_HOME>/mineru-venv-py310}"
LD_SO="$GLIBC_ROOT/lib64/ld-linux-x86-64.so.2"
PYTHON="$VENV/bin/python3.10"
ALL_LIBS="$GLIBC_ROOT/lib64:$GCC_ROOT/lib64:\
$DTK_ROOT/.hyhal/rocm_smi/lib:$DTK_ROOT/lib:$DTK_ROOT/lib64:\
$DTK_ROOT/hip/lib:$DTK_ROOT/dcc/lib:$DTK_ROOT/dcc/gcvm/lib:\
/usr/local/lib64:/usr/lib64"
exec "$LD_SO" --library-path "$ALL_LIBS" "$PYTHON" "$@"
Save as <RUN_DIR>/dcu-python and run chmod +x. Preflight checks:
<RUN_DIR>/dcu-python - <<'PY'
import ctypes, platform, sys
import onnxruntime, torch
print("python", sys.version.split()[0])
print("libc", platform.libc_ver())
print("torch", torch.__version__)
print("onnxruntime", onnxruntime.__version__)
print("devices", torch.cuda.device_count())
print("arch", torch.cuda.get_device_properties(0).gcnArchName)
PY
14.5 Local model configuration
Do not rely on a default filename in the old working directory. Explicitly set a configuration file for the reproduction:
{
"models-dir": {
"pipeline": "<PIPELINE_MODEL_ROOT>",
"vlm": "<VLM_MODEL_ROOT>"
},
"model-source": "local",
"device-mode": "cuda"
}
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>
export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1
If the configuration path is missing, model initialization later reports AttributeError: 'NoneType' object has no attribute 'get'. This does not indicate a corrupt model or a DCU operator failure.
14.6 Dynamic-library setup for spawned workers
MinerU 3.4.4 uses a PDF-render executor spawn. The reproduction script must set the executable before creating the process pool:
import multiprocessing
LAUNCHER = "<RUN_DIR>/dcu-python"
multiprocessing.set_executable(LAUNCHER)
try:
multiprocessing.get_context("spawn").set_executable(LAUNCHER)
except Exception:
pass
In production, put the same logic in sitecustomize.py. For a one-off reproduction, set it explicitly in the script to avoid modifying the shared virtual environment. In either case, verify that the launcher path exists.
14.7 Call the pipeline API directly
This revalidation does not use mineru.cli.client: the old CLI starts a local API child process. If that child uses the unwrapped sys.executable, it falls back to the system glibc 2.17.
import multiprocessing
import sys, time
from pathlib import Path
LAUNCHER = "<RUN_DIR>/dcu-python"
multiprocessing.set_executable(LAUNCHER)
try:
multiprocessing.get_context("spawn").set_executable(LAUNCHER)
except Exception:
pass
from mineru.cli.common import do_parse, read_fn
def main():
pdf = Path(sys.argv[1])
output = Path(sys.argv[2])
data = read_fn(str(pdf))
started = time.perf_counter()
do_parse(
output_dir=str(output),
pdf_file_names=[pdf.stem],
pdf_bytes_list=[data],
p_lang_list=["ch"],
backend="pipeline",
)
print(f"PARSE_ELAPSED={time.perf_counter() - started:.3f}")
if __name__ == "__main__":
main()
spawn reimports the main module; without a if __name__ == "__main__":guard, the child process reruns the entire parsing script.
14.8 Minimal Slurm job
#!/bin/bash -l
#SBATCH -p <DCU_PARTITION>
#SBATCH --gres=dcu:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=27gb
#SBATCH --time=00:20:00
#SBATCH -J mineru-reproduce
#SBATCH -o <LOG_DIR>/mineru-reproduce-%j.out
#SBATCH -e <LOG_DIR>/mineru-reproduce-%j.err
set -u
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VENV>/bin/activate
export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>
export OMP_NUM_THREADS=8
export ORT_NUM_THREADS=8
<RUN_DIR>/dcu-python \
<RUN_DIR>/reproduce.py \
<INPUT_PDF> \
<OUTPUT_DIR>
rc=$?
echo "PARSE_EXIT_RC=$rc"
exit "$rc"
After submission, check both the scheduler state and the workload exit code:
sbatch reproduce.slurm
sacct -j <JOB_ID> -X -o State,ExitCode,Elapsed -P -n
grep -E "PARSE_ELAPSED|PARSE_EXIT_RC" <LOG_FILE>
# COMPLETED | 0:0
# PARSE_EXIT_RC=0
14.9 Result and artifact verification
| Item | 2026-09-27 revalidation |
|---|---|
| Input | 14-page technical-paper PDF |
| Pipeline exit code | 0 |
| Cold-start parsing time | 142.9 s |
| Model initialization | 17.2 s |
| Inference during the processing window | 45.6 s, 0.307 pages/s |
| FP32 / FP16 / BF16 | 9.123 / 16.147 / 5.764 TFLOPS |
| Markdown | 30,483 bytes |
middle.json | 893,576 bytes |
| Images extracted | 27 files |
test -s <OUTPUT_DIR>/*/auto/*.md
test -s <OUTPUT_DIR>/*/auto/*_middle.json
test -s <OUTPUT_DIR>/*/auto/*_layout.pdf
find <OUTPUT_DIR> -type f | sort
grep -c '^#' <OUTPUT_DIR>/*/auto/*.md
grep -o '!\[\](images/' <OUTPUT_DIR>/*/auto/*.md | wc -l
An exit code of 0 is only the first acceptance check. Also verify that the Markdown is nonempty, the structured JSON parses, images were written to disk, and headings, equations, and captions appear in the output.
14.10 Errors and diagnosis
| Symptom | Root cause | Resolution |
|---|---|---|
ld-linux... No such file / rc 127 | Wrapper points to an archived glibc directory | Resolve the current runtime paths; do not modify the system ELF |
CXXABI_1.3.8 not found | glibc/DTK loaded, but GCC runtime is missing | Load the GCC 12.2 module and add its lib64 to the runtime closure |
GLIBC_2.27 not found | API child process started by the CLI bypasses the loader | Call do_parse()directly or repair the service-entry wrapper |
BrokenProcessPool | Spawned worker uses the system Python | multiprocessing.set_executable() |
NoneType ... get | Model configuration path has moved | Set the path explicitly and check it before startup MINERU_TOOLS_CONFIG_JSON |
pthread_setaffinity Repeated warning output | ONNX Runtime thread affinity conflicts with the cgroup CPU set | Limit the thread count; this did not affect correctness in this test |
torchvision.io zlib warning from | Optional image extension resolves to the system zlib | The pipeline succeeded; validate this extension separately before using it |
These steps show that the MinerU 3.4.4 pipeline can be restored in this Z100/DTK 26.04 environment. They do not establish compatibility for every CLI or service entry point, VLM engine, or future release. After upgrading a wheel, module, or model, repeat the audit starting at §14.2.
Appendix: Command reference
15.1 Environment
# Connect
ssh -p <SSH_PORT> <CLUSTER_HOST>
# Run these commands first in the job
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <USER_HOME>/mineru-venv-py310/bin/activate
# Offline mode (compute nodes have no external network access)
export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>
# Use GCC 12.2.0; GCC 11.2.0 and 13.3.0 lack libisl.so.15
export CC=<SOFTWARE_ROOT>/compiler/gcc-12.2.0/bin/gcc
# Limit threads when workers share a node
export OMP_NUM_THREADS=4 ORT_NUM_THREADS=4 OPENBLAS_NUM_THREADS=4 MKL_NUM_THREADS=4
# On CentOS 7, launch Python through the §14 wrapper;
# bare python or sys.executable falls back to system glibc 2.17.
15.2 Job template
#!/bin/bash
#SBATCH -p kshdnormal
#SBATCH --gres=dcu:4 # QOS requires at least 1
#SBATCH --cpus-per-task=32
#SBATCH --mem=111gb # limit = CPUs × 3569 MB
#SBATCH --time=02:00:00
#SBATCH -o job-%j.out
#SBATCH -e job-%j.err
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <USER_HOME>/mineru-venv-py310/bin/activate
...
rc=$?; echo "EXIT_RC=$rc"; exit $rc # Required so sacct reports the actual exit status
Check the account limits without entering the queue first:
sbatch --test-only job.slurm
# "Job N to start at ..." → request is within quota
# "Requested node configuration is not available" → no free nodes (quota is valid)
# "too much memory" → request exceeds the memory limit
15.3 Service pool
~/mineru_pool_ctl.sh status # jobs, live workers, and gateway health
~/mineru_pool_ctl.sh scale 2 # keep two pool jobs (four cards each)
~/mineru_pool_ctl.sh stop # stop jobs and clear the registry
~/mineru_pool_ctl.sh endpoints # print gateway addresses
15.4 Load testing
python3 loadtest.py
python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 12 4
15.5 Troubleshooting
# Do not rely on sacct State alone; inspect the job log
sacct -j <ID> -X -o State,Elapsed,ExitCode -P -n
# Filter verbose ONNX Runtime affinity messages
grep -av "pthread_setaffinity" job-<ID>.err | tail -30
# Check whether the process is working or stalled
ssh "cat /proc//status | grep -E '^State|^Threads'"
ssh "ls /proc//fd | grep -c kfd" # >0 means a DCU handle is open
# Check the wheel glibc requirements (use version sort: sort -V)
objdump -T .so | grep -oE "GLIBC_[0-9.]+" | sort -u -V | tail -3
15.6 Retagging a manylinux wheel for installation
When the runtime already uses glibc 2.28, pip's platform check is overly restrictive:
cp pkg-1.0-cp310-cp310-manylinux_2_28_x86_64.whl \
pkg-1.0-cp310-cp310-manylinux2014_x86_64.whl
pip install --no-deps --force-reinstall pkg-1.0-cp310-cp310-manylinux2014_x86_64.whl
15.7 Hygon DCU software downloads
# Package index (torchvision is under vision/)
curl -sL "https://download.sourcefind.cn:65024/directlink/4/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vision/DAS1.8/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vllm/DAS1.8/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/triton/DAS1.8/"
Versions must match the PyTorch build exactly; for example: ...das.opt1.dtk2604.torch271.
16. Data and job index
| Data | Job ID | Source file |
|---|---|---|
| Hardware capabilities (initial probe) | <JOB_ID> | data/job-logs.txt |
| Hardware capabilities (with warm-up) | <JOB_ID> | Same test |
| First successful parse | <JOB_ID> | data/sample-parse-output.md |
| CPU single-document run | <JOB_ID> | data/job-logs.txt |
| VLM backend | <JOB_ID> | Same test |
| DCU batch throughput | <JOB_ID> | Same test |
| CPU batch throughput | <JOB_ID> | Same test |
| Hygon Triton | <JOB_ID> | Same test |
| vLLM kernel | <JOB_ID> | Same test |
| Service pool v1 | <JOB_ID> | data/loadtest.log |
| Service pool v2 | <JOB_ID> | data/loadtest2.log、data/soak.log |
| Environment recovery after drift | <JOB_ID> | Redacted module, loader, and configuration paths; error-chain records |
| Full revalidation on 2026-09-27 | <JOB_ID> | 14-page pipeline logs and inventory of Markdown/JSON/image artifacts |