| Case | System | Atoms | Ensemble | HIP 2026.3 (ns/day) | CUDA 2023.2 (ns/day) | Speedup |
|---|---|---|---|---|---|---|
| villin | small protein | ~40K | NVT | 323.49 | 101.18 | 3.20× |
| rnase_cubic | Soluble protein | 24,040 | NVT | 147.94 | 66.93 | 2.21× |
| waterbox8 | pure water | 50,859 | NPT | 67.64 | 41.66 | 1.62× |
| adh_dodec | membrane protein | 95,561 | NVT | 26.09 | 18.89 | 1.38× |
| aqp_ensemble | Aquaporin | ~80K | NVT | 13.58 | 10.62 | 1.28× |
| ion_channel | Ion channel | ~50K | NVT | 20.17 | 15.87 | 1.27× |
| stmv | Large virus capsid | ~Millions | NVT | 2.30 | 1.83 | 1.26× |
Method: gmx_bench.slurm ran each backend/system combination three times. The options -notunepme -resetstep N/2 keep the timed region consistent. Results are reported as mean ± standard deviation; every case had a standard deviation below 2%.
Smaller systems showed larger speedups. Their GPUs are less fully utilized, so the lower startup and scheduling overhead of the native HIP kernels has a greater effect.
The previous report measured 15.31 vs 41.22 ns/day for waterbox8 NPT, making HIP appear 2.7× slower. That result was the only evidence then suggesting lower HIP performance. The idle-node retest produced the following results:
| Backend | Contended node | Idle node |
|---|---|---|
| HIP 2026.3 | 15.31 (affected by competing loads) | 67.64 |
| CUDA 2023.2 | 41.22 | 41.66 |
Cause: CPU contention on the node substantially affected the HIP backend. On the same contended node, HIP throughput fell from 148 to 9.4 ns/day (15× slower), while the CUDA-translated build remained near 66.8 ns/day. The earlier waterbox8 HIP result was likewise measured under contention. The resulting hypotheses about slow waterbox8 NPT performance—including virial rollback, DPP reduction, and register occupancy—were therefore based on contaminated measurements. Section 4 reports the retests of those optimization hypotheses.
For each case, the principal energy terms in the Energies section of run.log (LJ SR, Coulomb SR, potential, and total energy) agreed within approximately 0.01–0.05%. Temperatures stayed near the 300 K target, pressure remained within the expected ensemble fluctuations, and no NaN or Inf values appeared. The small differences are consistent with trajectory divergence caused by floating-point operations being accumulated in a different order; they do not indicate a physical error.
100 ps waterbox8 NPT run (50,000 steps): total energy −5.718 × 10⁵ kJ/mol, temperature 299.88 K, pressure −1.8 bar, no drift or abnormal termination, and throughput of 67.5 ns/day, consistent with the benchmark.
Architecture-probe mismatch (gfx90a vs gfx906): gmxManageHipccConfig.cmake used --offload-arch=gfx90a to test whether hipcc supported several optimization flags. DTK 26.04 does not provide a gfx90a device library, so every probe failed and all optimization flags were silently discarded. Performance fell to about 17 ns/day, with no reported error. The fix was to probe gfx906, add -mcpu=gfx906, and remove the unsupported -fno-gpu-rdc and -fno-slp-vectorize checks. This was the build change with the largest performance impact.
| Hypothesis | Test | Finding |
|---|---|---|
deviceStream.synchronize() accounts for most synchronization overhead | Remove the synchronization point and rebuild a control | No change; not the root cause |
CPU fallback for virial force reduction (useGpuFBufferOps falls back to CPU in computeVirial) | Confirm the fallback in the source; compare on an exclusive node | The fallback exists, but the controlled test did not identify it as a major bottleneck |
Dynamic DPP lane permutation followed by ds_bpermute causes 12.8× LDS traffic | Replace it with an atomic-reduction control | No change; reduction is not the bottleneck |
Raising minBlocksPerMp from 8 to 14 should reduce VGPR use | Measure VGPR count | VGPR use rose from 32 to 36. __launch_bounds__ did not behave as expected with this hipcc version; the roughly 1% gain was reverted. |
squeue to check for other jobs or request --exclusive; ② run each case three times and rerun it if the standard deviation exceeds 2% (this check is built into gmx_bench.slurm); ③ for cross-node comparisons, run both backends on the same set of exclusive nodes. Toolchain: DTK 26.04 (HIP 6.3.26113 / Clang 17), GCC 12.2.0, CMake 3.28.6, and OpenMPI 4.1.5.
# 0. Restore cmake_minimum_required to 3.28 (revert any downgrade to 3.25)
cd $SRC
sed -i 's/cmake_minimum_required(VERSION 3\.25)/cmake_minimum_required(VERSION 3.28)/g' CMakeLists.txt
find . -name CMakeLists.txt -exec sed -i 's/cmake_minimum_required *(VERSION 3\.25)/cmake_minimum_required(VERSION 3.28)/g' {} \;
# 1. gfx906 architecture-probe patch (critical; see §4.1)
cd $SRC/cmake
sed -i 's/--offload-arch=gfx90a/--offload-arch=gfx906/g' gmxManageHipccConfig.cmake
sed -i 's/gmx_hip_check_single_flag("-fno-gpu-rdc")/# REMOVED for gfx906: gmx_hip_check_single_flag("-fno-gpu-rdc")/' gmxManageHipccConfig.cmake
sed -i 's/gmx_hip_check_single_flag("-fno-slp-vectorize")/# REMOVED for gfx906: gmx_hip_check_single_flag("-fno-slp-vectorize")/' gmxManageHipccConfig.cmake
sed -i '/gmx_hip_check_single_flag("-ffast-math")/a\ gmx_hip_check_single_flag("-mcpu=gfx906")' gmxManageHipccConfig.cmake
# 2. Environment
module purge
module load mpi/openmpi/gcc-9.3.0/4.1.5
source <DTK_ROOT>/env.sh
export PATH=<GCC_ROOT>/bin:$PATH
export LD_LIBRARY_PATH=<GCC_ROOT>/lib64:$LD_LIBRARY_PATH
export amd_comgr_DIR=<DTK_ROOT>/lib64/cmake/amd_comgr
# 3. HIP Clang wrapper (adds AMD GPU optimization flags)
cat > $BUILD/hip_wrapper.sh << 'EOF'
#!/bin/bash
exec <DTK_ROOT>/llvm/bin/clang++ \
--gcc-toolchain=<GCC_ROOT> \
-mllvm -amdgpu-early-inline-all=true \
-mllvm -amdgpu-function-calls=false \
"$@"
EOF
chmod +x $BUILD/hip_wrapper.sh
# 4. Configure and build with CMake
$CMAKE $SRC \
-DGMX_BUILD_OWN_FFTW=ON -DGMX_GPU=HIP -DGMX_MPI=ON \
-DCMAKE_C_COMPILER=$GCC_HOME/bin/gcc -DCMAKE_CXX_COMPILER=$GCC_HOME/bin/g++ \
-DCMAKE_HIP_COMPILER=$BUILD/hip_wrapper.sh \
-DCMAKE_PREFIX_PATH=<DTK_ROOT> \
-Damd_comgr_DIR=<DTK_ROOT>/lib64/cmake/amd_comgr \
-DGMX_SIMD=AVX2_256 -DGMX_GPU_FFT_LIBRARY=rocFFT \
-DGMX_GPU_NB_DISABLE_CLUSTER_PAIR_SPLIT=ON \
-DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_X=1 -DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_Y=1 -DGMX_GPU_NB_NUM_CLUSTER_PER_CELL_Z=1 \
-DGMX_HIP_TARGET_ARCH=gfx906 -DAMDGPU_TARGETS=gfx906 -DGPU_TARGETS=gfx906 \
-DCMAKE_BUILD_TYPE=Release
make -j8
Full build script: <USER_HOME>/scripts/gromacs_cmake28_build.slurm. Submit it with sbatch to reproduce the build.
The validated build is located at <USER_HOME>/softwares/gromacs/build_cmake28/:
| file | Size | Description |
|---|---|---|
bin/gmx_mpi | 113 KB | Main executable (dynamically linked) |
lib/libgromacs_mpi.so.11.0.0 | 45 MB | Compute library containing the HIP kernels |
Run the build with a DTK runtime compatible with the one used to compile it:
module purge
module load mpi/openmpi/gcc-9.3.0/4.1.5
source <DTK_ROOT>/env.sh
export LD_LIBRARY_PATH=<GCC_ROOT>/lib64:$LD_LIBRARY_PATH
/path/to/build_cmake28/bin/gmx_mpi mdrun -s case.tpr -nb gpu -pme gpu ...
GROMACS 2026.3 writes version 138 .tpr files, which the cluster's prebuilt 2023.2 cannot read. Version 2026.3 can read older .tpr files. For cross-version comparisons, use input files generated by the corresponding GROMACS version.