opensource_community

GROMACS 2026.3 Port and Benchmark on Hygon Z100

Test supercomputing cluster · Hygon Z100 DCU (gfx906) · Final report
Test date: 2026-08-16  |  Operator: <USER>  |  Comparison baseline: GROMACS 2023.2 (CUDA translation)

1. Summary

Cases with a speedup
7 / 7
HIP build 2026.3 vs CUDA-translated build 2023.2
Maximum speedup
3.20×
villin (small protein NVT)
Minimum speedup
1.26×
stmv (Million Atom Virus NVT)
waterbox8 (previously anomalous case)
1.62×
67.64 vs 41.66 ns/day, reversed from "2.7× slower"
✅ Acceptance decision:The new version meets the “no regression” criterion and is faster in all seven tested cases. The earlier result showing waterbox8 NPT as 2.7× slower was caused by CPU contention on the test node (see Section 3).

2. Performance results (idle node; three runs per case)

CaseSystemAtomsEnsembleHIP 2026.3
(ns/day)
CUDA 2023.2
(ns/day)
Speedup
villinsmall protein~40KNVT323.49101.183.20×
rnase_cubicSoluble protein24,040NVT147.9466.932.21×
waterbox8pure water50,859NPT67.6441.661.62×
adh_dodecmembrane protein95,561NVT26.0918.891.38×
aqp_ensembleAquaporin~80KNVT13.5810.621.28×
ion_channelIon channel~50KNVT20.1715.871.27×
stmvLarge virus capsid~MillionsNVT2.301.831.26×

Method: gmx_bench.slurm ran each backend/system combination three times. The options -notunepme -resetstep N/2 keep the timed region consistent. Results are reported as mean ± standard deviation; every case had a standard deviation below 2%. Smaller systems showed larger speedups. Their GPUs are less fully utilized, so the lower startup and scheduling overhead of the native HIP kernels has a greater effect.

3. Why the waterbox8 result changed

The previous report measured 15.31 vs 41.22 ns/day for waterbox8 NPT, making HIP appear 2.7× slower. That result was the only evidence then suggesting lower HIP performance. The idle-node retest produced the following results:

BackendContended nodeIdle node
HIP 2026.315.31 (affected by competing loads)67.64
CUDA 2023.241.2241.66

Cause: CPU contention on the node substantially affected the HIP backend. On the same contended node, HIP throughput fell from 148 to 9.4 ns/day (15× slower), while the CUDA-translated build remained near 66.8 ns/day. The earlier waterbox8 HIP result was likewise measured under contention. The resulting hypotheses about slow waterbox8 NPT performance—including virial rollback, DPP reduction, and register occupancy—were therefore based on contaminated measurements. Section 4 reports the retests of those optimization hypotheses.

3. Correctness verification

For each case, the principal energy terms in the Energies section of run.log (LJ SR, Coulomb SR, potential, and total energy) agreed within approximately 0.01–0.05%. Temperatures stayed near the 300 K target, pressure remained within the expected ensemble fluctuations, and no NaN or Inf values appeared. The small differences are consistent with trajectory divergence caused by floating-point operations being accumulated in a different order; they do not indicate a physical error.

100 ps waterbox8 NPT run (50,000 steps): total energy −5.718 × 10⁵ kJ/mol, temperature 299.88 K, pressure −1.8 bar, no drift or abnormal termination, and throughput of 67.5 ns/day, consistent with the benchmark.

4. Build issues and performance analysis

4.1 Compiler configuration failure and its effect

Architecture-probe mismatch (gfx90a vs gfx906): gmxManageHipccConfig.cmake used --offload-arch=gfx90a to test whether hipcc supported several optimization flags. DTK 26.04 does not provide a gfx90a device library, so every probe failed and all optimization flags were silently discarded. Performance fell to about 17 ns/day, with no reported error. The fix was to probe gfx906, add -mcpu=gfx906, and remove the unsupported -fno-gpu-rdc and -fno-slp-vectorize checks. This was the build change with the largest performance impact.

4.2 Performance optimization assumptions and verification results

HypothesisTestFinding
deviceStream.synchronize() accounts for most synchronization overheadRemove the synchronization point and rebuild a controlNo change; not the root cause
CPU fallback for virial force reduction (useGpuFBufferOps falls back to CPU in computeVirial)Confirm the fallback in the source; compare on an exclusive nodeThe fallback exists, but the controlled test did not identify it as a major bottleneck
Dynamic DPP lane permutation followed by ds_bpermute causes 12.8× LDS trafficReplace it with an atomic-reduction controlNo change; reduction is not the bottleneck
Raising minBlocksPerMp from 8 to 14 should reduce VGPR useMeasure VGPR countVGPR use rose from 32 to 36. __launch_bounds__ did not behave as expected with this hipcc version; the roughly 1% gain was reverted.
Benchmarking principle:Establish a reproducible baseline before tuning. Resource contention can distort measurements and lead to incorrect conclusions about kernel-level optimizations.

5. Benchmark preconditions

Node-state checks:The HIP backend was more sensitive to CPU contention, with up to a 4× performance difference in this test; the CUDA-translated build was more stable. Before benchmarking: ① use squeue to check for other jobs or request --exclusive; ② run each case three times and rerun it if the standard deviation exceeds 2% (this check is built into gmx_bench.slurm); ③ for cross-node comparisons, run both backends on the same set of exclusive nodes.

6. Build procedure and reproduction steps

Toolchain: DTK 26.04 (HIP 6.3.26113 / Clang 17), GCC 12.2.0, CMake 3.28.6, and OpenMPI 4.1.5.

# 0. Restore cmake_minimum_required to 3.28 (revert any downgrade to 3.25)
cd $SRC
sed -i 's/cmake_minimum_required(VERSION 3\.25)/cmake_minimum_required(VERSION 3.28)/g' CMakeLists.txt
find . -name CMakeLists.txt -exec sed -i 's/cmake_minimum_required *(VERSION 3\.25)/cmake_minimum_required(VERSION 3.28)/g' {} \;

# 1. gfx906 architecture-probe patch (critical; see §4.1)
cd $SRC/cmake
sed -i 's/--offload-arch=gfx90a/--offload-arch=gfx906/g' gmxManageHipccConfig.cmake
sed -i 's/gmx_hip_check_single_flag("-fno-gpu-rdc")/# REMOVED for gfx906: gmx_hip_check_single_flag("-fno-gpu-rdc")/' gmxManageHipccConfig.cmake
sed -i 's/gmx_hip_check_single_flag("-fno-slp-vectorize")/# REMOVED for gfx906: gmx_hip_check_single_flag("-fno-slp-vectorize")/' gmxManageHipccConfig.cmake
sed -i '/gmx_hip_check_single_flag("-ffast-math")/a\    gmx_hip_check_single_flag("-mcpu=gfx906")' gmxManageHipccConfig.cmake

# 2. Environment
module purge
module load mpi/openmpi/gcc-9.3.0/4.1.5
source <DTK_ROOT>/env.sh
export PATH=<GCC_ROOT>/bin:$PATH
export LD_LIBRARY_PATH=<GCC_ROOT>/lib64:$LD_LIBRARY_PATH
export amd_comgr_DIR=<DTK_ROOT>/lib64/cmake/amd_comgr

# 3. HIP Clang wrapper (adds AMD GPU optimization flags)
cat > $BUILD/hip_wrapper.sh << 'EOF'
#!/bin/bash
exec <DTK_ROOT>/llvm/bin/clang++ \
  --gcc-toolchain=<GCC_ROOT> \
  -mllvm -amdgpu-early-inline-all=true \
  -mllvm -amdgpu-function-calls=false \
  "$@"
EOF
chmod +x $BUILD/hip_wrapper.sh

# 4. Configure and build with CMake
$CMAKE $SRC \
  -DGMX_BUILD_OWN_FFTW=ON -DGMX_GPU=HIP -DGMX_MPI=ON \
  -DCMAKE_C_COMPILER=$GCC_HOME/bin/gcc -DCMAKE_CXX_COMPILER=$GCC_HOME/bin/g++ \
  -DCMAKE_HIP_COMPILER=$BUILD/hip_wrapper.sh \
  -DCMAKE_PREFIX_PATH=<DTK_ROOT> \
  -Damd_comgr_DIR=<DTK_ROOT>/lib64/cmake/amd_comgr \
  -DGMX_SIMD=AVX2_256 -DGMX_GPU_FFT_LIBRARY=rocFFT \
  -DGMX_GPU_NB_DISABLE_CLUSTER_PAIR_SPLIT=ON \
  -DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_X=1 -DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_Y=1 -DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_Z=1 \
  -DGMX_HIP_TARGET_ARCH=gfx906 -DAMDGPU_TARGETS=gfx906 -DGPU_TARGETS=gfx906 \
  -DCMAKE_BUILD_TYPE=Release
make -j8

Full build script: <USER_HOME>/scripts/gromacs_cmake28_build.slurm. Submit it with sbatch to reproduce the build.

7. Build artifacts

The validated build is located at <USER_HOME>/softwares/gromacs/build_cmake28/:

fileSizeDescription
bin/gmx_mpi113 KBMain executable (dynamically linked)
lib/libgromacs_mpi.so.11.0.045 MBCompute library containing the HIP kernels

Run the build with a DTK runtime compatible with the one used to compile it:

module purge
module load mpi/openmpi/gcc-9.3.0/4.1.5
source <DTK_ROOT>/env.sh
export LD_LIBRARY_PATH=<GCC_ROOT>/lib64:$LD_LIBRARY_PATH

/path/to/build_cmake28/bin/gmx_mpi mdrun -s case.tpr -nb gpu -pme gpu ...

GROMACS 2026.3 writes version 138 .tpr files, which the cluster's prebuilt 2023.2 cannot read. Version 2026.3 can read older .tpr files. For cross-version comparisons, use input files generated by the corresponding GROMACS version.

8. Feature changes (GROMACS 2023.2—2026.3)