opensource_community
Full Experiment Archive · Redacted Public Edition

MinerU on Hygon Z100 DCU

This archive documents environment validation, service-pool implementation, and performance testing. It records the results, reproduction commands, measurements, and source job IDs for each approach, along with the evidence used to test the underlying assumptions.

Cluster Test cluster (national supercomputing facility) Hardware 4 × Hygon Z100 DCUs (gfx906) Operating system CentOS 7.6 / glibc 2.17 Software stack DTK 26.04 Test dates 2026-08-12 to 2026-08-13; revalidated 2026-09-27 Deployment model Bare metal (systemd)
01

Key findings

MetricResultEvidence
Service-pool throughput955 docs/hour(4 cards, concurrency 12)soak.log
Stability72 requests; 0 failures; 0 rejections§10
Fault recoveryWorker killed during load test; restarted automatically without losing requests§10.3
Memory leakNone (no RSS growth across 72 requests)§10.4
Backend selectionpipeline 3.9 s/page; VLM at 14.93 s/page is not recommended§8.1
FP16 / FP32 / BF1616.15 / 9.01 / 5.83 TFLOPS§3
Triton / torch.compile✕ Architecture limitation Not supported on gfx906§4
vLLM HIP kernel✓ Available PagedAttention verified in testing§5
CPU queueEquivalent output; 8.9× slower at steady state§8.3
Reproduction after environment drift✓ Passed 14-page pipeline run; exit code 0§14
Summary

MinerU can run in production on Z100 using the pipeline backend. Triton is unavailable on this architecture, but MinerU does not require it.

02

Background and validation scope

2.1 Initial environment status

The prior work set up the environment and produced a V2 report claiming that “all tests passed.” A review found one critical gap:

ssh -p <SSH_PORT> <CLUSTER_HOST> 'ls ~/.cache/modelscope/hub ~/models ~/output 2>/dev/null'
# all directories are empty

MinerU had never parsed a document. The V2 report's “end-to-end test” measured only magika file-type detection—the first preprocessing step in MinerU, not document parsing. No models had been downloaded.

2.2 Reusable software and data

ItemPath / version
venv<USER_HOME>/mineru-venv-py310(Python 3.10.14)
torch2.7.1+das.opt1.dtk2604
MinerUVersion 3.4.4 (editable install; source patched for magika compatibility)
glibc 2.28<USER_HOME>/glibc-2.28-rpm/ (path used at the time; relocated during revalidation, see §14)
Wrappersdcu-python / mineru / mineru-api

2.3 Cluster resources and runtime constraints

cat ~/<SKILL_CONFIG>
ConstraintValue
Memory limitcpus-per-task × 3569 MB
DCU queuekshdnormal, QOS enforces --gres=dcu:1
CPU queuekshcnormal, several nodes; no GRES restriction
External network on compute nodes❌ None
/tmpNot shared across nodes
03

Hardware capabilities and software environment

3.1 Initial probe results and limitations

Probe job <JOB_ID>:

sbatch ~/probe_caps.slurm     # see cluster/probe_caps.py
torch 2.7.1 devices 4
arch: gfx906:sramecc-:xnack- mem GB: 17.2
[OK]   bf16 matmul: supported=True maxerr=0.250
[OK]   fp16 matmul: 4.85 TFLOPS          ← suspicious: slower than FP32
[OK]   fp32 matmul: 8.91 TFLOPS
[FAIL] triton kernel: FileNotFoundError: '.../gcc-11.2.0-install/bin/gcc'
[OK]   SDPA: flash_avail=True mem_eff=True
[FAIL] torch.compile: FileNotFoundError: (same error)
[OK]   torch._int_mm: ok (512, 512)
[OK]   onnxruntime: 1.16.3 providers=['AzureExecutionProvider','CPUExecutionProvider']

FP16 performance below FP32 was unexpected—the investigation found that the benchmark had no warm-up.

Benchmark method after warm-up was added

Job <JOB_ID>, with five warm-up iterations followed by 30 timed iterations:

def bench(dt, n=2048, iters=30):
    a = torch.randn(n, n, device="cuda", dtype=dt)
    b = torch.randn(n, n, device="cuda", dtype=dt)
    for _ in range(5): c = a @ b          # ← warm-up
    torch.cuda.synchronize(); s = time.time()
    for _ in range(iters): c = a @ b
    torch.cuda.synchronize()
    return 2 * n**3 / ((time.time() - s) / iters) / 1e12
fp32: 9.01 TFLOPS
fp16: 16.15 TFLOPS      ← 1.8× FP32
bf16: 5.73 TFLOPS
--- force the rocBLAS/hipBLAS backend ---
fp16: 16.12 TFLOPS      ← unchanged; rocBLAS was already in use
bf16: 5.83 TFLOPS

3.3 Final capability matrix

CapabilityResultNotes
FP16 GEMM16.15 TFLOPSUse FP16 for inference
FP32 GEMM9.01 TFLOPS
BF16 GEMM5.83 TFLOPSSoftware emulation; slowest option
torch._int_mm✓INT8 available
SDPA⚠ DegradedFalls back to the math backend; no Flash or memory-efficient attention
Multi-GPU✓ 4 cards16 GB per card (the report states 17.2 GB)
hipBLASLt✕ No gfx906 kernelThe library contains kernels only for gfx928; runtime warnings are emitted
FP8 / MFMA✕Vega20 architecture limitation
Methodological lesson

Benchmarks involving one-time costs (kernel tuning, JIT compilation, or model loading) must include warm-up; otherwise, the conclusion may be completely reversed.

04

Triton build failure analysis

This case shows why environment failures must be traced across layers: the initial error did not identify the underlying cause.

4.1 Incorrect GCC path

FileNotFoundError: '<USER_HOME>/software/gcc-11.2.0-install/bin/gcc'

This directory appeared in the user's PATH, but it did not exist:

echo $PATH | tr ':' '\n' | grep gcc
# <USER_HOME>/software/gcc-11.2.0-install/bin   ← invalid path
ls <USER_HOME>/software/gcc-11.2.0-install/bin/gcc
# ls: cannot access ...: No such file or directory

Triton initially appeared unavailable; investigation showed that the configured GCC path pointed to a nonexistent directory.

4.2 System GCC version and library-path issue

After correcting the path to the existing gcc-11.2.0:

cc1: error while loading shared libraries: libisl.so.15: cannot open shared object file

Compilers installed on the cluster, tested individually:

echo "int main(){return 0;}" > /tmp/t.c
for cc in <SOFTWARE_ROOT>/compiler/gcc-11.2.0/bin/gcc \
          <SOFTWARE_ROOT>/compiler/gcc-12.2.0/bin/gcc \
          <SOFTWARE_ROOT>/compiler/gcc-13.3.0/bin/gcc \
          /opt/rh/devtoolset-7/root/usr/bin/gcc \
          <SOFTWARE_ROOT>/compiler/rocm/dtk-26.04/llvm/bin/clang; do
  $cc /tmp/t.c -o /tmp/t.out 2>/dev/null && echo "OK   $cc" || echo "FAIL $cc"
done
CompilerResult
gcc-11.2.0✕ Missing libisl.so.15
gcc-12.2.0✓
gcc-13.3.0✕ Missing libisl.so.15
devtoolset-7✓
DTK clang✓

4.3 Root cause: unsupported gfx906 target

After switching to GCC 12.2.0 (job <JOB_ID>):

loc("probe_triton.py":4:0): error: unsupported target: 'gfx906'
RuntimeError: PassManager::run failed

4.4 Validation against Hygon's Triton build

Hygon also publishes a DCU-specific Triton build, which warranted a separate test:

curl -sL "https://download.sourcefind.cn:65024/directlink/4/triton/DAS1.8/"
# triton-3.1.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl

Install it in an isolated directory to avoid affecting the working environment:

cp triton-3.1.0+...manylinux_2_28_x86_64.whl triton-3.1.0+...manylinux2014_x86_64.whl
pip install --no-deps --target=<USER_HOME>/dcu-triton-test 

The first run could not find libgcvm.so.17git. The library was in dtk-26.04/dcc/gcvm/lib, which was missing from the wrapper's library path. Adding that directory to dcu-python's ALL_LIBS resolved the issue (job <JOB_ID>):

triton: 3.1.0 from <USER_HOME>/dcu-triton-test/triton/__init__.py
arch: gfx906:sramecc-:xnack-
[FAIL] RuntimeError: PassManager::run failed
error: unsupported target: 'gfx906'      ← same result as the PyPI build

Source inspection confirmed the architecture allowlist:

grep -ohE "gfx9[0-9]{2}[a-z]*" triton/backends/amd/compiler.py | sort -u
# gfx928 gfx936 gfx940 gfx941 gfx942
grep -ohE "gfx9[0-9]{2}[a-z]*" triton/backends/hcu/compiler.py | sort -u
# gfx928 gfx936 gfx938 gfx940 gfx941 gfx942

gfx906 is not listed.

⚠ A misleading clue

strings libtriton.so | grep gfx906 A search does find gfx906. That entry is in LLVM upstream's AMDGPU target list. LLVM recognizing an architecture does not mean Triton supports it. Test the installed build; do not infer support from strings in the binary.

Conclusion

Neither Triton 3.x (the PyPI build nor Hygon's DCU build) supports gfx906,torch.compile so Triton-dependent features are unavailable This is an architectural limitation; configuration changes cannot resolve it.

05

vLLM porting feasibility

5.1 Backend capability boundaries

First, distinguish two easily confused terms:

  • VLM = Vision-Language Model, a type ofmodel(MinerU2.5-Pro-2605-1.2B is one example)
  • vLLM is an inference engine that provides PagedAttention.

MinerU has two layers of options, differing by one letter:

Parsing backend (how the document is parsed) ├── pipeline Specialized small models: layout + OCR + formulas + tables ├── vlm-engine A large VLM reads each page end to end └── hybrid-engine └─ Selects an inference engine within vlm-engine ├── transformers PyTorch only; slowest ├── vllm-engine ← vLLM is used here └── lmdeploy

5.2 gfx906 kernel architecture validation

Hygon also provides a DCU build of vLLM:

curl -sL "https://download.sourcefind.cn:65024/directlink/4/vllm/DAS1.8/"
# vllm-0.11.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl

Inspect the packaged code objects:

strings vllm/_C.abi3.so     | grep -oE "gfx[0-9]{3,4}" | sort -u
# gfx906 gfx926 gfx928 gfx936 gfx938
strings vllm/_moe_C.abi3.so | grep -oE "gfx[0-9]{3,4}" | sort -u
# gfx906 gfx926 gfx928 gfx936 gfx938

Hygon excludes gfx906 from Triton's supported targets but retains it in vLLM's HIP kernels.

5.3 Runtime test (job <JOB_ID>)

from vllm import _custom_ops as ops
# PagedAttention is vLLM's core attention operation
ks = torch.tensor(1.0, device="cuda"); vs = torch.tensor(1.0, device="cuda")
ops.paged_attention_v1(out, q, kc, vc, nkvh, 1.0/(hd**0.5),
                       btab, slen, bs, bs*2, None, "auto", ks, vs, 0,0,0,64,0)
arch: gfx906:sramecc-:xnack-
[OK] import vllm 0.11.0
[OK] vllm._C loaded
[OK] rms_norm (HIP): maxerr=0.0010
[OK] rotary_embedding (HIP): ran
[OK] paged_attention_v1 (HIP) ★: out(4, 8, 64) finite=True

5.4 Conclusion and remaining work

vLLM has a working execution path on Z100—all core HIP kernels work, the opposite of the Triton result.

Unresolved: platform detection fails (is_rocm: False). vLLM uses import amdsmi to detect ROCm:

# vllm/platforms/__init__.py:109
import amdsmi
amdsmi.amdsmi_init()

but DTK 26.04 provides only the older rocm_smi interface (.hyhal/rocm_smi/bin/rsmiBindings.py). Completing inference requires an amdsmi shim or a patch to rocm_platform_plugin().

No benefit for MinerU: vLLM's strengths are PagedAttention and continuous batching for high-concurrency, long-form generation. MinerU's VLM workload uses a single-page image and short output, so it does not benefit from these features. The vLLM path is still worth pursuing for an LLM inference service on Z100.

06

Parsing workflow failures and fixes

Each new error showed that diagnosis had progressed to the next layer. The issues below are listed in the order encountered:

6.1 Model not available

modelscope imports torch and fails on the login node, which has no DCU driver. Use huggingface_hub (pure Python) with hf-mirror instead:

# Check endpoint reachability
curl -s -o /dev/null -w "%{http_code}" https://hf-mirror.com/     # 200
curl -s -o /dev/null -w "%{http_code}" https://huggingface.co/    # blocked from this network
# cluster/dl_models.py
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
os.environ["HF_HOME"] = "<USER_HOME>/mineru-models/hf"
from huggingface_hub import snapshot_download
for repo in ["opendatalab/MinerU2.5-Pro-2605-1.2B",
             "opendatalab/PDF-Extract-Kit-1.0"]:
    snapshot_download(repo, max_workers=8)

This downloads 20 GB. Set the local path:

{
  "models-dir": {
    "pipeline": "<USER_HOME>/mineru-models/models/OpenDataLab--PDF-Extract-Kit-1.0/snapshots/master",
    "vlm":      "<USER_HOME>/mineru-models/models/OpenDataLab--MinerU2.5-Pro-2605-1.2B/snapshots/master"
  },
  "model-source": "local",
  "device-mode": "cuda"
}

6.2 Resolve libssl.so.3 not found

File "mineru/cli/fast_api.py", line 19, in 
    import uvicorn
ImportError: libssl.so.3: cannot open shared object file
Error: Local mineru-api exited before becoming healthy.

Root cause: The MinerU CLI launches a mineru-api child process with the system sys.executable, bypassing the dcu-python wrapper.

Workaround used at the time: Call the do_parse() Python API instead of the CLI. Permanent service fix: Add /usr/local/lib64 to the wrapper's ALL_LIBS.

6.3 operator torchvision::nms does not exist

torchvision 0.21.0 from PyPI is incompatible with the DCU Torch 2.7.1 ABI.

The DCU torchvision build is published in the sourcefind vision/ directory, not torchvision/:

curl -sL "https://download.sourcefind.cn:65024/directlink/4/vision/DAS1.8/"
# torchvision-0.22.0+das.opt1.dtk2604.torch271-cp310-...whl   ← built for torch 2.7.1

Because the system glibc is 2.17, pip rejects the manylinux_2_28 wheel. Retagging lets pip install it; the program itself already runs under the user-space glibc 2.28 loader:

W=torchvision-0.22.0+das.opt1.dtk2604.torch271-cp310-cp310-manylinux_2_28_x86_64.whl
cp $W ${W/manylinux_2_28/manylinux2014}
pip install --no-deps --force-reinstall ${W/manylinux_2_28/manylinux2014}

6.4 Missing OCR/layout dependencies

pip install shapely pyclipper omegaconf einops ftfy

6.5 Resolve BrokenProcessPool

concurrent.futures.process.BrokenProcessPool:
  A process in the process pool was terminated abruptly

Root cause: MinerU renders PDFs using a spawn process pool:

# mineru/utils/pdf_image_tools.py:160
if start_method != "spawn":
    return ProcessPoolExecutor(max_workers=max_workers,
                               mp_context=multiprocessing.get_context("spawn"))

The spawn method starts a fresh interpreter through sys.executable, which loses the configured library paths.

Fix——sitecustomize.py, imported automatically at Python startup so the change applies globally:

# /lib/python3.10/site-packages/sitecustomize.py
import os
_w = "<USER_HOME>/mineru-venv-py310/bin/dcu-python"
if os.path.exists(_w):
    import multiprocessing
    multiprocessing.set_executable(_w)
    try:
        multiprocessing.get_context("spawn").set_executable(_w)
    except Exception:
        pass

Verification:

dcu-python -c "
import multiprocessing as mp
from concurrent.futures import ProcessPoolExecutor
def f(x): return x*x
if __name__=='__main__':
    with ProcessPoolExecutor(2, mp_context=mp.get_context('spawn')) as ex:
        print(list(ex.map(f,[1,2,3])))"
# [1, 4, 9]
⚠ Caller warning

Any script that calls MinerU must include an if __name__ == "__main__": guard. A spawned worker imports the main module; without the guard, the script runs again, producing duplicate [info] output and then a BrokenProcessPool error. The MinerU documentation does not mention this requirement.

6.6 Resolve Unsupported model IR version: 10

onnxruntime.capi.onnxruntime_pybind11_state.Fail:
  Load model from .../PP-LCNet_x1_0_table_cls.onnx failed:
  Unsupported model IR version: 10, max supported IR version: 9

onnxruntime 1.16.3 supports IR up to v9, but MinerU's table-classification model uses IR v10.

The V2 report concluded that ONNX Runtime ≥1.17 was incompatible because it was published only for manylinux_2_28. The wheel can in fact be retagged using the method in §6.3; this step is required.

curl -sL "https://pypi.tuna.tsinghua.edu.cn/simple/onnxruntime/" \
  | grep -oE 'href="[^"]*onnxruntime-1\.20\.1-cp310-cp310-manylinux_2_27_x86_64[^"]*"'
curl -sL -o ort.whl "https://pypi.tuna.tsinghua.edu.cn/packages/63/47/.../onnxruntime-1.20.1-...whl"
cp ort.whl onnxruntime-1.20.1-cp310-cp310-manylinux2014_x86_64.whl
pip install --no-deps --force-reinstall onnxruntime-1.20.1-cp310-cp310-manylinux2014_x86_64.whl

6.7 Additional fixes

IssueFix
onnxruntime emits pthread_setaffinity an errorSet OMP/ORT/OPENBLAS/MKL_NUM_THREADS=4
libgcvm.so.17git is missingAdd dtk-26.04/dcc/gcvm/lib to the ALL_LIBS
No module named mineru.cli.apiIn version 3.4.4 the module is named fast_api; update the mineru-api wrapper
07

Output quality validation

Job <JOB_ID>; 23-page Mooncake paper, 0.58 MB.

dcu-python parse_api.py ~/sourcecode/Mooncake/Mooncake-v3.pdf ~/parse-out/api-pipeline pipeline
[RESULT] backend=pipeline elapsed=146.8s
[OUT] .../Mooncake-v3.md chars=81072
processing-window multi-file infer finished, cost: 90.52, speed: 0.254 page/s

Output files:

Mooncake-v3.md                  81 KB    Markdown source
Mooncake-v3_content_list.json  121 KB    structured content
Mooncake-v3_middle.json        2.1 MB    intermediate representation
Mooncake-v3_layout.pdf         835 KB    layout visualization
Mooncake-v3_span.pdf           832 KB    span visualization
images/                        26 files

Checks performed (not just relying on a zero job exit code):

ItemResultVerification
Markdown81,072 charactersHeading hierarchy, superscript author symbols ♠♡, and abstract structure all correct
Tables3 HTML tablesSampled LRUCache/LFUCache rows; values and rows are complete
Images23 references / 26 filesimages/ Files present on disk
Headings31Hierarchy matches the source
Inline equations17 LaTeX expressions$T_{queue}$、$\mathrm{MLP}$、$T_{prefill}$
Display equationsNoneThe source uses inline math; no missed display equations

Excerpt from the first page:

# Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Ruoyu Qin♠♡1 Zheming Li♠1 ... ♠Moonshot AI ♡Tsinghua University

## Abstract

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI...
Lesson from verification

Searching only for $$ display equations returned no matches and nearly led to the wrong conclusion that formula recognition had failed. The paper uses inline math: middle.json contains 38 inline_equation entries, and MFR identifies 19 regions. Choose checks that match the content; the absence of one format does not prove a feature is missing.

08

Performance tests and results

8.1 Backend comparison

BackendJobThroughputAssessment
pipeline (DCU)<JOB_ID>3.9 s/page(90.5 s / 23 pages)Selected for production
pipeline (CPU)<JOB_ID>10.1 s/page (231.3 s)Equivalent output; see §8.3
VLM (transformers)<JOB_ID>14.93 s/pageTimed out after 50 minutes; incomplete

VLM is 3.8× slower and has no acceleration path on gfx906. Triton is unavailable, so the vllm-engine path cannot run. VLM is not recommended for production.

Confirming that VLM was computing rather than stalled:

ssh <COMPUTE_NODE> "cat /proc/<PID>/status | grep -E '^State|^Threads'; ls /proc/<PID>/fd | grep -c kfd"
# State: R (running)   Threads: 44   kfd handles: 1

8.2 Model loading dominates runtime

Share of runtime spent loading the model:

146.8s
Single document (cold start)
22.5s
Same document within a batch
~85%
Model-loading share of runtime

Model loading accounts for about 85% of the runtime. This rules out the “one job per document” design: 85% of the allocated time would be spent reloading the same models. Keep workers resident and process documents in batches.

8.3 CPU-queue output equivalence

Byte-for-byte comparison of the output:

diff ~/parse-out/cpu/Mooncake-v3/auto/Mooncake-v3.md \
     ~/parse-out/api-pipeline/Mooncake-v3/auto/Mooncake-v3.md
# Five-line diff:
# 98,99c98
# < ## Listing 1: Request samples.
# ---
# > Listing 1: Request samples.

81,074 vs 81,072 characters; the only difference is one heading level. The CPU queue is a viable production path, not merely a fallback.

However, a single-document comparison substantially understates the performance gap:

Single documentBatch steady state
DCU146.8 s19.6 s
CPU315.4 s173.8 s
Ratio2.1×8.9×

For one document, model loading dominates both paths, making CPU and DCU runtimes appear similar. This understates the steady-state throughput gap by about four times.

On the CPU path, formula recognition is the main cost: MFR runs at 5.86 it/s on DCU versus 3.85 s/it on CPU (about 22× slower).

8.4 Batch throughput (without HTTP)

DCU job <JOB_ID>; CPU job <JOB_ID>:

dcu-python throughput.py ~/pdf-corpus ~/parse-out/tput pipeline 4
ConfigurationIncluding cold startSteady stateSteady-state time per document
4 DCU233 docs/h735 docs/h19.6 s
4 CPU worker63 docs/h83 docs/h173.8 s

With load balanced across four workers, runtimes were 176.4 / 177.4 / 185.3 / 185.7 s, a spread of less than 6%.

Bug found: workers did not exit after batch processing and remained stuck, each of the four processes retaining 3.8 GB RSS. MinerU's _get_pdf_render_executor() creates a persistentpool; its non-daemon child processes prevent the parent from exiting. Fix:

from mineru.utils.pdf_image_tools import shutdown_pdf_render_executor
shutdown_pdf_render_executor()   # called during worker shutdown

After the fix, the CPU batch job exits normally and prints BENCH_DONE and RC=0 exits cleanly.

8.5 Queue configuration comparison

Cluster load during the measurements:

sinfo -p kshdnormal -o "%t %D" -h | sort    # alloc 735 / idle 2  (fully loaded)
sinfo -p kshcnormal -o "%t %D" -h | sort    # idle 1106           (idle)

Total throughput = throughput per node × number of available nodes:

  • A few idle DCU nodes × 735 ≈ 1,470 docs/h
  • Several idle CPU nodes × 83 ≈ 91,663 docs/h

The order-of-magnitude difference comes from node availability, not compute capacity. Use the CPU queue for large offline batches; reserve DCUs for latency-sensitive online requests and formula-heavy documents.

09

Service-pool architecture and implementation

9.1 Network reachability and service architecture

Job <JOB_ID>. First verify whether compute nodes can serve requests:

# Start the service on a compute node and write its endpoint
IP=$(hostname -I | awk '{print $1}'); PORT=$((30000 + RANDOM % 5000))
dcu-python svc_hello.py $PORT &
echo "$IP:$PORT" > ~/svc_endpoint.txt

# Access it from the login node
curl --max-time 8 "http://$(cat ~/svc_endpoint.txt)/"
# MINERU_SVC_ALIVE on <COMPUTE_NODE>      ← reachable
TestResult
Compute node binds to 0.0.0.0 Self-test✓
Login node → compute-node port✓
Compute node → external network✕ (expected)

Cluster-internal network connectivityis available; this is separate from the lack of external access on compute nodes. The service pool can therefore expose a reachable service endpoint.

9.2 Service-pool architecture

┌───────────────────────────────┐ Client ──HTTP──►│ Gateway :38080 │ │ Health checks · least-load scheduling │ │ Per-worker limits · retry on failure │ └───────────┬───────────────────┘ │ Dispatch to an idle slot ┌──────────┬─────────┴─────────┬──────────┐ ▼ ▼ ▼ ▼ worker :38100 :38101 :38102 :38103 DCU 0 DCU 1 DCU 2 DCU 3 └──────────┴─────────┬─────────┴──────────┘ ▼ Registry (shared directory; file mtime is heartbeat)

Design rationale(Each decision is based on measurements, not assumptions):

DecisionRationale
One process per cardMinerU models are not thread-safe; concurrent requests in one process serialize
HIP_VISIBLE_DEVICES Pin each worker to a cardIsolates GPU memory; an OOM in one worker does not affect the others
Limit each worker to one concurrent requestEstablished by load testing; see §10.1
Use file modification time as a heartbeatWorkers register and deregister themselves; the gateway scans the directory, with no separate service-discovery component
Retry failed requests on another workerParsing is idempotent, so retrying is safe

9.3 MinerU API endpoints

grep -n "@app\.\(get\|post\)" -A3 mineru/cli/fast_api.py
RoutePurpose
POST /file_parseSynchronous parsing
POST /tasksAsynchronous submission
GET /tasks/{id} /tasks/{id}/resultAsynchronous status query
GET /healthHealth check (for gateway)

9.4 Key implementation excerpts

Gateway scheduling (scripts/mineru_gateway.py):

def _acquire(exclude=()):
    """Reserve a slot on the least-loaded healthy worker, waiting if all busy."""
    deadline = time.time() + SLOT_WAIT_TIMEOUT
    with _cv:
        while True:
            cands = [(v["inflight"], v["served"], u)
                     for u, v in _workers.items()
                     if v["healthy"] and u not in exclude
                     and v["inflight"] < MAX_INFLIGHT]      # rate limit
            if cands:
                cands.sort()
                url = cands[0][2]
                _workers[url]["inflight"] += 1
                return url
            if time.time() >= deadline: return None
            _cv.wait(2.0)                                    # queue instead of rejecting

Retry after failure:

for attempt in range(MAX_ATTEMPTS):
    url = _acquire(exclude=tuple(tried))     # exclude workers that have already failed
    ...
    except urllib.error.HTTPError as e:
        if e.code < 500:                     # client error (4xx): do not retry
            return 
        _release(url, False)
    tried.append(url)

Worker supervision (cluster/mineru_worker.sh):

while true; do
  mineru-api --host 0.0.0.0 --port $PORT &
  APID=$!
  # Register only after health checks pass; do not route traffic during model loading
  for i in $(seq 1 90); do
    curl -sf --max-time 3 "http://127.0.0.1:$PORT/health" >/dev/null && READY=1 && break
    sleep 2
  done
  [ "$READY" = 1 ] && echo "{\"url\":\"$URL\",...}" > "$REGFILE"
  # Heartbeat: touch while healthy; the gateway checks the modification time
  while kill -0 $APID <PID>>/dev/null; do
    curl -sf --max-time 5 ".../health" >/dev/null && touch "$REGFILE"
    sleep 10
  done
  rm -f "$REGFILE"; sleep 5      # restart after a crash
done

9.5 Startup

sbatch ~/mineru_pool.slurm        # cluster/mineru_pool.slurm
[t+95s] healthy workers: 4
=== POOL READY ===
{"pool_size": 4, "healthy": 4, "max_inflight_per_worker": 1, ...}

Verified from the login node:

curl -s "http://<SERVICE_ENDPOINT>/pool/status" | python3 -m json.tool
10

Service-pool performance and stability tests

10.1 Performance comparison between gateway versions

Same workload: 24 requests at concurrency 8

python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 8 2
Metricv1 (no concurrency limit or retry)v2 (production version)
Error rate8.3% (2 × 502)0.0%
Throughput374 docs/h440 docs/h (+18%)
Load distribution11 / 4 / 5 / 26 / 5 / 7 / 6
p5027.2 s38.0 s
p90111.9 s121.7 s

Raw output:data/loadtest.log、data/loadtest2.log

Concurrency limits improved throughput—MinerU models are not thread-safe, so requests within a worker are processed serially. Sending more work to a busy worker only creates an internal queue. Load skew fell from 5.5:1 to 1.4:1.

The increase in p50 latency is expected: v1's lower p50 was an artifact of a few requests reaching idle workers; p90 latency was higher and 8.3% of requests failed. v2 queues requests fairly across the pool.

10.2 Overload test at three times worker capacity

48 requests at concurrency 12:

python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 12 4
total=48  ok=48  failed=0  error_rate=0.0%
wall=180.9s  throughput=955 docs/hour  concurrency=12
latency  p50=32.4s  p90=70.4s  p99=131.1s  min=11.9s  max=131.1s
per-worker distribution:  11 / 11 / 14 / 12

Pool status during the load test:

"in_flight": 4  "waiting_for_slot": 8  "queue_peak": 8  "retries": 0  "rejected": 0

Backpressure worked as designed: four requests were allowed in flight; the other eight queued instead of failing.The pool slowed under overload without dropping requests.

Throughput rose from 440 to 955 docs/hour as concurrency increased from 8 to 12, showing that the pipeline was not saturated at concurrency 8. The 955 docs/hour resultexceeded the offline batch rate of 735because static per-card sharding leaves workers idle, while the service pool dispatches work dynamically to whichever slot becomes available.

10.3 Worker recovery under load

PID=$(python3 -c "import json;print(json.load(open('.../<COMPUTE_NODE>-dcu3.json'))['pid'])")
ssh <COMPUTE_NODE> "kill -9 $PID"        # PID <PID>

After 50 seconds:

before: "healthy": 3  "total_served": 5   "total_failed": 0
after:  "healthy": 4  "total_served": 16  "total_failed": 0

Confirmed worker replacement by PID change and logs:

cat ~/mineru-registry/<COMPUTE_NODE>-dcu3.json
# {"url":"http://<SERVICE_ENDPOINT>",...,"pid":<PID>}     ← 原 <PID>

grep -a dcu3 pool-*.err | grep -E "restart|READY|exited"
# [worker dcu3] exited, restarting in 5s (total restarts: 1)
# [worker dcu3] starting on http://<SERVICE_ENDPOINT> (restart #1)
# [worker dcu3] READY -> registered http://<SERVICE_ENDPOINT>

Recovery was tested under real load;total_failed there were zero failures and no lost requests.

10.4 Memory stability across 72 requests

ssh <COMPUTE_NODE> "ps -u <USER> -o rss,cmd --sort=-rss | grep -a fast_api | awk '{print \$1/1024}'"
w0w1w2w3
Before4839478047544744
After4837477347624742

No growth. The process count remained at four, with no buildup of spawned children.

10.5 Summary

72
Total requests (24 + 48)
0
Failures
0
Rejected
955/h
Peak throughput
11

Approaches not adopted and operating limits

Failed paths and their costs are documented to avoid repeating the same work.

11.1 Embedding library paths in the ELF with patchelf ✕ Failed

Rationale: The ld-linux wrapper covers only the top-level process. Embedding library paths in the Python binary with patchelf might let every child process inherit them, which would be cleaner in theory.

patchelf --set-interpreter $GLIBC_DIR/ld-linux-x86-64.so.2 python3.10
Inconsistency detected by ld.so: dl-call-libc-early-init.c: 37:
  _dl_call_libc_early_init: Assertion `sym != NULL' failed!

Even with an added rpath, the process still segfaulted; ldd also crashed:

patchelf --set-interpreter ... --force-rpath --set-rpath "$G:/usr/local/lib64:/usr/lib64" python3.10
./python3.10 -c "print(1)"        # Segmentation fault
ldd python3.10                    # Segmentation fault

Isolation tests:

TestResult
Change only the rpath; keep the system interpreter✓ Passed
Change only the interpreter✕ Assertion failed
Change both✕ segfault
Conclusion

The glibc 2.28 ld.so loader cannot be combined with the system glibc 2.17 libraries, so the change was reverted from backup. Do not patch the ELF interpreter across this glibc version boundary; use a wrapper and sitecustomize.py.

11.2 VLM backend ⚠ Runs, but is not practical

At 14.93 s/page, the 23-page run did not finish within the 50-minute limit. No acceleration path is available on gfx906: Triton is unavailable, so the vllm-engine / lmdeploy engine paths cannot run; only single-sequence inference through Transformers is available. The pipeline backend already produces higher-quality output and is 3.8× faster.

11.3 Full vLLM inference ⏸ Incomplete, not failed

Core HIP kernels work in testing (§5), but platform detection is blocked at import amdsmi. DTK 26.04 provides only the older rocm_smi. A shim or patch is needed.The path is viable; the integration work is unfinished.

11.4 Failed approaches recorded in the V2 report

ApproachResultReason
Install onnxruntime with Conda✕Channel access timed out
Singularity container✕Docker Hub was blocked
Python 3.8✕MinerU requires Python ≥3.10
Force-install Python 3.12 + onnx 1.17✕ (at the time)This work resolved the same class of issue by retagging the wheel
12

Assumptions tested and revised

Three course corrections during the work changed the direction of the conclusions. They are recorded here because the method matters as much as the results.

12.1“FP16 offers no benefit” ← benchmark had no warm-up

Without warm-up, FP16 measured 4.85 TFLOPS, below FP32 at 8.91 TFLOPS, suggesting that FP16 offered no benefit on gfx906. After warm-up, FP16 reached 16.15 TFLOPS, 1.8× the FP32 result.

Lesson

Warm up any benchmark that includes one-time costs.

12.2“vLLM cannot run on gfx906” ← inferred from Triton's failure

The initial assessment tied vLLM availability on gfx906 to Triton support, without distinguishing their dependency layers. vLLM's PagedAttention uses hand-written HIP kernels and does not depend on Triton. Inspection confirmed that Hygon's _C.abi3.so contains kernels compiled for gfx906, and the paged_attention_v1 operation ran successfully.

Lesson

Do not infer one component's support from another component's failure, even when they appear related. Validate each layer separately.

12.3“CPU is only 2.1× slower” ← extrapolated from a single document

For one document, 315.4 s vs 146.8 s suggested a 2.1× difference, leading to the claim that “the CPU queue is only slightly slower.” Batch steady-state measurements were 173.8 s vs 19.6 s, an 8.9× difference.

A single-document run is dominated by model loading on both systems. Since model-loading times are similar on CPU and DCU, they mask the true throughput gap by a factor of four.

Lesson

Compare steady-state results when startup costs are fixed overhead.

12.4“Formula recognition is broken” ← checked only one output format

The initial check searched only for display equations delimited by $$ and found none. The paper instead uses inline math throughout: 17 $...$ expressions appear in the rendered text, and middle.json contains 38 inline_equation regions.

Lesson

Choose checks that match the content; the absence of one format does not prove a feature is missing.

12.5“ONNX Runtime no longer works” ← child process bypassed the loader

During environment revalidation on 2026-09-27, the top-level Python process used the user-space glibc 2.28 loader, but the local API process launched by the MinerU CLI used the system glibc 2.17 and reported GLIBC_2.27 not found. The wheel was valid; the runtime environment had not been propagated across the process boundary.

After loading the GCC 12.2 and DTK 26.04 modules, starting the main process with the same loader, and making multiprocessing children inherit that entry point, the 14-page pipeline completed and returned 0.

Lesson

Check library compatibility across the full process tree. A successful import in the parent does not prove that the CLI, service process, and spawned workers use the same runtime.

13

Deployment and runtime configuration

The deployment is bare metal, outside Slurm.

13.1 Runtime compatibility and operating-system requirements

The cluster's ld-linux wrapper was created to work around CentOS 7's glibc 2.17, but it leaves child-process blind spots (§6.2 and §6.5).With bare-metal deployment, the operating system is under your control, so this issue can be avoided.

The DCU Torch build requires GLIBC_2.25:

objdump -T torch.libs/libevent_core*.so* | grep -oE "GLIBC_[0-9.]+" | sort -u -V | tail -3
# GLIBC_2.14
# GLIBC_2.17
# GLIBC_2.25      ← actual requirement
⚠

Use sort -V for version-aware sorting. A lexicographic sort can place GLIBC_2.25 before GLIBC_2.9, leading to the false conclusion that 2.9 is the newest required version.

The Torch wheels tagged manylinux_2_17 and manylinux_2_28 have the same MD5 hash, 7b098c2f9c036fcdb9ae244228bac822. Only their tags differ; retagging does not remove the actual glibc requirement. Use Rocky Linux 8+, Ubuntu 20.04+, or CentOS 8+ to avoid the need for glibc wrappers and wheel retagging.

13.2 Deployment

./mineru-baremetal-deploy.sh preflight    # Check glibc / DTK / virtualenv / models
sudo ./mineru-baremetal-deploy.sh install # Generate the systemd unit
sudo ./mineru-baremetal-deploy.sh start
./mineru-baremetal-deploy.sh status

systemd handles startup and restarts the service after a crash with Restart=always. Set TimeoutStartSec=300 to allow time for model loading. The registration script registers a worker only after its health checks pass, so traffic is not sent to a worker that is still loading the model.

13.3 Requests

curl -X POST http://:8000/file_parse \
     -F "files=@paper.pdf" -F "backend=pipeline"

curl http://:8000/pool/status

The X-MinerU-Worker response header identifies the worker that handled the request; X-MinerU-Attempt records the retry attempt.

13.4 Monitoring

MetricAlert condition
healthy< Target threshold: alert
waiting_for_slotSustained > more than 2× the worker count → scale out
total_failedAlert on abnormal growth
retriesA sudden increase indicates an unstable worker
rejectedAny value above 0 requires immediate investigation

13.5 Scaling out

Scale horizontally without changing the gateway: deploy workers on new machines (without a gateway), point the registry to the same shared-storage path (NFS/GPFS), and the gateway will discover them automatically.

The maximum workers on one node depends on its card count. If CPU work (such as layout analysis and OCR preprocessing) becomes the bottleneck, reduce *_NUM_THREADS per-worker concurrency or the worker count.

14

Environment recovery and reproduction

On 2026-09-27, after the original experiment directory had been archived and relocated, the runtime closure was audited from scratch and an end-to-end pipeline parse was completed on a 14-page PDF.

Conclusion

The Kunshan module environmentdoes not provide a separate glibc module. Reproduction requires the GCC 12.2 and DTK 26.04 modules, a user-space glibc 2.28 loader, the correct local model configuration, and the same Python entry point for the parent and multiprocessing children.

14.1 Reproduction scope and evidence

LayerValidationEvidence
Login nodeNo standalone glibc module; GCC 12.2 and DTK 26.04 are availablemodule avail/show
Compute nodePyTorch import, device enumeration, and FP32/FP16/BF16 GEMMControlled single-card probe
Process treeMain Python process and PDF-render worker use the same loaderMultiprocessing test
End to end14-page PDF → Markdown / JSON / images / rendered PDFPipeline exit code 0

The public edition omits usernames, cluster addresses, node names, actual job IDs, and real paths. It retains versions, error types, command structure, performance figures, and acceptance criteria.

14.2 Environment audit

Do not rerun old commands blindly. First check whether the wrapper, configuration, and model paths still exist:

module -t avail 2>&1 | grep -Ei '(^|/)glibc([/-]|$)|compat-glibc'
# Expected: no output; glibc 2.28 is not provided by a module.

for f in \
  <VENV>/bin/dcu-python \
  <VENV>/bin/mineru \
  <VENV>/bin/mineru-api \
  <VENV>/lib/python3.10/site-packages/sitecustomize.py
do
  echo "=== $f ==="
  sed -n '1,80p' "$f"
done

test -x <GLIBC_ROOT>/lib64/ld-linux-x86-64.so.2
test -x <VENV>/bin/python3.10
test -f <MODEL_CONFIG>
test -d <PIPELINE_MODEL_ROOT>

This revalidation found that the four old entry points still referenced an archived glibc directory. The model weights were still present, but the configuration file had moved to a new scripts directory.The environment may still be intact even when its entry-point paths are stale.

14.3 Module environment

module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VENV>/bin/activate

module list
# compiler/gcc/12.2.0
# compiler/dtk/26.04

gcc --version | head -1
# gcc (GCC) 12.2.0

echo "$ROCM_PATH"
# <SOFTWARE_ROOT>/compiler/rocm/dtk-26.04

The DTK module supplies HIP and math libraries, while the GCC module supplies a newer libstdc++.so.6. Neither replaces glibc. Loading only DTK, or adding only glibc to --library-path, can still result in CXXABI_1.3.8 not found.

14.4 A single Python entry point

Avoid maintaining several slightly different wrappers. Create one entry point and use it for the CLI, service, and spawned workers:

#!/bin/bash
set -u

GLIBC_ROOT="${GLIBC_ROOT:-<USER_HOME>/softwares/runtime/glibc-2.28}"
GCC_ROOT="${GCC_ROOT:-<SOFTWARE_ROOT>/compiler/gcc-12.2.0}"
DTK_ROOT="${DTK_ROOT:-<SOFTWARE_ROOT>/compiler/rocm/dtk-26.04}"
VENV="${VENV:-<USER_HOME>/mineru-venv-py310}"

LD_SO="$GLIBC_ROOT/lib64/ld-linux-x86-64.so.2"
PYTHON="$VENV/bin/python3.10"
ALL_LIBS="$GLIBC_ROOT/lib64:$GCC_ROOT/lib64:\
$DTK_ROOT/.hyhal/rocm_smi/lib:$DTK_ROOT/lib:$DTK_ROOT/lib64:\
$DTK_ROOT/hip/lib:$DTK_ROOT/dcc/lib:$DTK_ROOT/dcc/gcvm/lib:\
/usr/local/lib64:/usr/lib64"

exec "$LD_SO" --library-path "$ALL_LIBS" "$PYTHON" "$@"

Save as <RUN_DIR>/dcu-python and run chmod +x. Preflight checks:

<RUN_DIR>/dcu-python - <<'PY'
import ctypes, platform, sys
import onnxruntime, torch
print("python", sys.version.split()[0])
print("libc", platform.libc_ver())
print("torch", torch.__version__)
print("onnxruntime", onnxruntime.__version__)
print("devices", torch.cuda.device_count())
print("arch", torch.cuda.get_device_properties(0).gcnArchName)
PY

14.5 Local model configuration

Do not rely on a default filename in the old working directory. Explicitly set a configuration file for the reproduction:

{
  "models-dir": {
    "pipeline": "<PIPELINE_MODEL_ROOT>",
    "vlm": "<VLM_MODEL_ROOT>"
  },
  "model-source": "local",
  "device-mode": "cuda"
}
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>
export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1

If the configuration path is missing, model initialization later reports AttributeError: 'NoneType' object has no attribute 'get'. This does not indicate a corrupt model or a DCU operator failure.

14.6 Dynamic-library setup for spawned workers

MinerU 3.4.4 uses a PDF-render executor spawn. The reproduction script must set the executable before creating the process pool:

import multiprocessing

LAUNCHER = "<RUN_DIR>/dcu-python"
multiprocessing.set_executable(LAUNCHER)
try:
    multiprocessing.get_context("spawn").set_executable(LAUNCHER)
except Exception:
    pass

In production, put the same logic in sitecustomize.py. For a one-off reproduction, set it explicitly in the script to avoid modifying the shared virtual environment. In either case, verify that the launcher path exists.

14.7 Call the pipeline API directly

This revalidation does not use mineru.cli.client: the old CLI starts a local API child process. If that child uses the unwrapped sys.executable, it falls back to the system glibc 2.17.

import multiprocessing
import sys, time
from pathlib import Path

LAUNCHER = "<RUN_DIR>/dcu-python"
multiprocessing.set_executable(LAUNCHER)
try:
    multiprocessing.get_context("spawn").set_executable(LAUNCHER)
except Exception:
    pass

from mineru.cli.common import do_parse, read_fn

def main():
    pdf = Path(sys.argv[1])
    output = Path(sys.argv[2])
    data = read_fn(str(pdf))
    started = time.perf_counter()
    do_parse(
        output_dir=str(output),
        pdf_file_names=[pdf.stem],
        pdf_bytes_list=[data],
        p_lang_list=["ch"],
        backend="pipeline",
    )
    print(f"PARSE_ELAPSED={time.perf_counter() - started:.3f}")

if __name__ == "__main__":
    main()
Why the main guard matters

spawn reimports the main module; without a if __name__ == "__main__":guard, the child process reruns the entire parsing script.

14.8 Minimal Slurm job

#!/bin/bash -l
#SBATCH -p <DCU_PARTITION>
#SBATCH --gres=dcu:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=27gb
#SBATCH --time=00:20:00
#SBATCH -J mineru-reproduce
#SBATCH -o <LOG_DIR>/mineru-reproduce-%j.out
#SBATCH -e <LOG_DIR>/mineru-reproduce-%j.err

set -u
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VENV>/bin/activate

export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>
export OMP_NUM_THREADS=8
export ORT_NUM_THREADS=8

<RUN_DIR>/dcu-python \
  <RUN_DIR>/reproduce.py \
  <INPUT_PDF> \
  <OUTPUT_DIR>
rc=$?
echo "PARSE_EXIT_RC=$rc"
exit "$rc"

After submission, check both the scheduler state and the workload exit code:

sbatch reproduce.slurm
sacct -j <JOB_ID> -X -o State,ExitCode,Elapsed -P -n
grep -E "PARSE_ELAPSED|PARSE_EXIT_RC" <LOG_FILE>
# COMPLETED | 0:0
# PARSE_EXIT_RC=0

14.9 Result and artifact verification

Item2026-09-27 revalidation
Input14-page technical-paper PDF
Pipeline exit code0
Cold-start parsing time142.9 s
Model initialization17.2 s
Inference during the processing window45.6 s, 0.307 pages/s
FP32 / FP16 / BF169.123 / 16.147 / 5.764 TFLOPS
Markdown30,483 bytes
middle.json893,576 bytes
Images extracted27 files
test -s <OUTPUT_DIR>/*/auto/*.md
test -s <OUTPUT_DIR>/*/auto/*_middle.json
test -s <OUTPUT_DIR>/*/auto/*_layout.pdf

find <OUTPUT_DIR> -type f | sort
grep -c '^#' <OUTPUT_DIR>/*/auto/*.md
grep -o '!\[\](images/' <OUTPUT_DIR>/*/auto/*.md | wc -l

An exit code of 0 is only the first acceptance check. Also verify that the Markdown is nonempty, the structured JSON parses, images were written to disk, and headings, equations, and captions appear in the output.

14.10 Errors and diagnosis

SymptomRoot causeResolution
ld-linux... No such file / rc 127Wrapper points to an archived glibc directoryResolve the current runtime paths; do not modify the system ELF
CXXABI_1.3.8 not foundglibc/DTK loaded, but GCC runtime is missingLoad the GCC 12.2 module and add its lib64 to the runtime closure
GLIBC_2.27 not foundAPI child process started by the CLI bypasses the loaderCall do_parse()directly or repair the service-entry wrapper
BrokenProcessPoolSpawned worker uses the system Pythonmultiprocessing.set_executable()
NoneType ... getModel configuration path has movedSet the path explicitly and check it before startup MINERU_TOOLS_CONFIG_JSON
pthread_setaffinity Repeated warning outputONNX Runtime thread affinity conflicts with the cgroup CPU setLimit the thread count; this did not affect correctness in this test
torchvision.io zlib warning fromOptional image extension resolves to the system zlibThe pipeline succeeded; validate this extension separately before using it
Scope of applicability

These steps show that the MinerU 3.4.4 pipeline can be restored in this Z100/DTK 26.04 environment. They do not establish compatibility for every CLI or service entry point, VLM engine, or future release. After upgrading a wheel, module, or model, repeat the audit starting at §14.2.

15

Appendix: Command reference

15.1 Environment

# Connect
ssh -p <SSH_PORT> <CLUSTER_HOST>

# Run these commands first in the job
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <USER_HOME>/mineru-venv-py310/bin/activate

# Offline mode (compute nodes have no external network access)
export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1
export MINERU_MODEL_SOURCE=local
export MINERU_DEVICE_MODE=cuda
export MINERU_TOOLS_CONFIG_JSON=<MODEL_CONFIG>

# Use GCC 12.2.0; GCC 11.2.0 and 13.3.0 lack libisl.so.15
export CC=<SOFTWARE_ROOT>/compiler/gcc-12.2.0/bin/gcc

# Limit threads when workers share a node
export OMP_NUM_THREADS=4 ORT_NUM_THREADS=4 OPENBLAS_NUM_THREADS=4 MKL_NUM_THREADS=4

# On CentOS 7, launch Python through the §14 wrapper;
# bare python or sys.executable falls back to system glibc 2.17.

15.2 Job template

#!/bin/bash
#SBATCH -p kshdnormal
#SBATCH --gres=dcu:4              # QOS requires at least 1
#SBATCH --cpus-per-task=32
#SBATCH --mem=111gb               # limit = CPUs × 3569 MB
#SBATCH --time=02:00:00
#SBATCH -o job-%j.out
#SBATCH -e job-%j.err
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <USER_HOME>/mineru-venv-py310/bin/activate
...
rc=$?; echo "EXIT_RC=$rc"; exit $rc     # Required so sacct reports the actual exit status

Check the account limits without entering the queue first:

sbatch --test-only job.slurm
# "Job N to start at ..."  → request is within quota
# "Requested node configuration is not available" → no free nodes (quota is valid)
# "too much memory" → request exceeds the memory limit

15.3 Service pool

~/mineru_pool_ctl.sh status       # jobs, live workers, and gateway health
~/mineru_pool_ctl.sh scale 2      # keep two pool jobs (four cards each)
~/mineru_pool_ctl.sh stop         # stop jobs and clear the registry
~/mineru_pool_ctl.sh endpoints    # print gateway addresses

15.4 Load testing

python3 loadtest.py    
python3 loadtest.py <SERVICE_ENDPOINT> ~/pdf-corpus 12 4

15.5 Troubleshooting

# Do not rely on sacct State alone; inspect the job log
sacct -j <ID> -X -o State,Elapsed,ExitCode -P -n

# Filter verbose ONNX Runtime affinity messages
grep -av "pthread_setaffinity" job-<ID>.err | tail -30

# Check whether the process is working or stalled
ssh  "cat /proc//status | grep -E '^State|^Threads'"
ssh  "ls /proc//fd | grep -c kfd"     # >0 means a DCU handle is open

# Check the wheel glibc requirements (use version sort: sort -V)
objdump -T .so | grep -oE "GLIBC_[0-9.]+" | sort -u -V | tail -3

15.6 Retagging a manylinux wheel for installation

When the runtime already uses glibc 2.28, pip's platform check is overly restrictive:

cp pkg-1.0-cp310-cp310-manylinux_2_28_x86_64.whl \
   pkg-1.0-cp310-cp310-manylinux2014_x86_64.whl
pip install --no-deps --force-reinstall pkg-1.0-cp310-cp310-manylinux2014_x86_64.whl

15.7 Hygon DCU software downloads

# Package index (torchvision is under vision/)
curl -sL "https://download.sourcefind.cn:65024/directlink/4/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vision/DAS1.8/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/vllm/DAS1.8/"
curl -sL "https://download.sourcefind.cn:65024/directlink/4/triton/DAS1.8/"

Versions must match the PyTorch build exactly; for example: ...das.opt1.dtk2604.torch271.

·

16. Data and job index

DataJob IDSource file
Hardware capabilities (initial probe)<JOB_ID>data/job-logs.txt
Hardware capabilities (with warm-up)<JOB_ID>Same test
First successful parse<JOB_ID>data/sample-parse-output.md
CPU single-document run<JOB_ID>data/job-logs.txt
VLM backend<JOB_ID>Same test
DCU batch throughput<JOB_ID>Same test
CPU batch throughput<JOB_ID>Same test
Hygon Triton<JOB_ID>Same test
vLLM kernel<JOB_ID>Same test
Service pool v1<JOB_ID>data/loadtest.log
Service pool v2<JOB_ID>data/loadtest2.log、data/soak.log
Environment recovery after drift<JOB_ID>Redacted module, loader, and configuration paths; error-chain records
Full revalidation on 2026-09-27<JOB_ID>14-page pipeline logs and inventory of Markdown/JSON/image artifacts
First edition: 2026-08-13 · Environment-recovery validation: 2026-09-27 · Data from redacted test records · Public report edition