opensource_community
← Return to test report

LAMMPS Hygon DCU build and performance validation on two platforms

gfx936/BW + gfx906/Z100 · DTK 25.04.2–26.04 · Kokkos/HIP · GPU package · GPU-aware MPI
Author: opensource_community · Consolidated update: 2026-08-18

1. Test platforms and verification scope

This report brings together two independent validations built around the same engineering workflow. Platform A covers gfx936 architecture support, GPU package comparisons, and source optimizations; Platform B covers native gfx906/Z100 builds, correctness, and multi-card scaling. Locations, cluster aliases, SSH addresses, partition and node names, and absolute user paths are redacted.

1.1 Platform A · gfx936/BW

SystemSugon 8.9 (Linux) CPUHygon C86 (128 cores, 64 cores/socket) DCUC-3000 BW, gfx936, 80 CU, 1500 MHz, 64 GB ToolchainGCC 9.3.0, DTK 25.04.2 / hipcc Clang 17 MPIMPICH 4.1.2; extended verification using OpenMPI 5.0.7 SoftwareCMake 3.24.1, LAMMPS 22 Jul 2025 Update 4, Kokkos 4.6.2

1.2 Platform B · gfx906/Z100

SystemCentOS 7.6 / glibc 2.17 CPUHygon C86 7185 (32 cores) DCUZ100 / Device 66a1 / Vega 20, gfx906, 64 CU, 16 GB/card, 4 cards per machine Node memory123 GB ToolchainGCC 9.3.0, DTK 26.04, hipcc → dcc 25.10.0-0 / Clang 17 SoftwareCMake 3.25.0, LAMMPS 22 Jul 2025 Update 5, Kokkos 4.6.2 MPIOpenMPI 5.0.7 (built from source with --with-rocm) SchedulerSlurm 3.2.10; CPU build and 1–4 DCU test partition names redacted

2. Build strategies and issue analysis

2.1 GPU package with the HIP backend

Configuration: -D PKG_GPU=ON -D GPU_API=HIP -D HIP_ARCH=gfx936

First result:Compiled successfully, but hung when running

Initial diagnosis:DTK's CMake configuration applies the -xhip flag to the GPU package, causing hipcc to compile every .cpp file for the device. The resulting executable hung during HIP initialization.

After the fix:Setting AMDGPU_TARGETS="gfx936" and disabling HIP_USE_DEVICE_SORT allowed the build to compile and run. Performance tests nevertheless showed the GPU package to be 4.6–6.3× slower than Kokkos (see the comparison in Section 6).

2.2 Kokkos HIP backend with the default architecture (mismatch)

Configuration: -D PKG_KOKKOS=ON -D Kokkos_ENABLE_HIP=ON

Result:The build succeeds, but execution fails with running kernels compiled for gfx906 on gfx936.

Analysis:Kokkos 4.6.2 does not recognize gfx936 and falls back to gfx906. Because the gfx906 and gfx936 instruction sets are incompatible, execution fails with hipErrorInvalidKernelFile.

2.3 Add gfx936 support to Kokkos

Approach:Add gfx936 support explicitly to Kokkos' CMake configuration and HIP backend.

📄 Modification 1: lib/kokkos/cmake/kokkos_arch.cmake
Register AMD_GFX936 and its gfx936 target name. List gfx936 before gfx906 so automatic detection selects the correct architecture first, and register the numeric alias 936.
📄 Modification 2: lib/kokkos/cmake/KokkosCore_config.h.in
Add #cmakedefine KOKKOS_ARCH_AMD_GFX936.
📄 Modification 3: lib/kokkos/core/src/HIP/Kokkos_HIP_Instance.hpp
Include defined(KOKKOS_ARCH_AMD_GFX936) in the HIPTraits condition that sets WarpSize. Hygon BW DCUs use a wavefront size of 64, as do gfx906 and gfx908.
📄 Modification 4: lib/kokkos/core/src/HIP/Kokkos_HIP_IsXnack.hpp
Add defined(KOKKOS_ARCH_AMD_GFX936) to the gpu_arch_can_access_system_allocations branch that returns false. Hygon DCUs do not support access to system allocations.

2.4 gfx906/Z100: native architecture path

Unlike gfx936, gfx906 (Vega 20) has been natively supported by Kokkos 4.6.2, so there is no need to modify the four architecture files of Kokkos and directly enable Kokkos_ARCH_AMD_GFX906=ON and AMDGPU_TARGETS=gfx906. This reduces architecture porting costs, but DTK 26.04 and the cluster toolchain expose another set of build issues.

2.5 DTK 26.04 cuda_wrappers host mode conflict

At about 98% of the build, main.cpp failed because cuda_wrappers/algorithm reported unknown type name '__host__'. In LAMMPS Update 5, cmake/CMakeLists.txt:854 forces the lmp target to -x c++ to address an nlohmann/json HIP SFINAE issue. DTK 26.04's hipcc still injects CUDA wrappers in host mode, causing the standard <algorithm> header to be parsed incorrectly.

📄 DTK 26.04 compatibility fix
Replace the target's -x c++ option with -xhip --offload-arch=gfx906:
target_compile_options(lmp PRIVATE
  "$<$<COMPILE_LANGUAGE:CXX>:SHELL:-xhip --offload-arch=gfx906>")
The same issue is not reproduced on Platform A in DTK 25.04.2 and is therefore logged as a DTK 26.04 regression rather than a general LAMMPS defect.

2.6 HSA library symlink unavailable on the CPU node

The libhsa-runtime64.so symlink resolves through the DCU driver directory /opt/hyhal/, which is absent from the CPU build node. CMake therefore cannot find the HSA runtime. Set HSA_RUNTIME_LIBRARY explicitly to the actual shared-library path; the public commands use the placeholder <DTK_ROOT>/hyhal/lib/libhsa-runtime64.so.

2.7 Broken make interpreter and login-node memory pressure

On Platform B, the system /usr/bin/make references a missing ELF interpreter at /opt/glibc-2.25/lib/ld-linux.so.2; load compiler/make/4.4 instead. The login node has about 251 GB of shared memory for roughly 249 users, with only about 46 GB available during testing. When hipcc compiled roughly 1,000 C++ files, dcc was killed by the OOM killer. The build was moved to the CPU compute queue and submitted with 32 cores and about 111 GB of exclusive memory.

2.8 Restrict the build to the target architecture

DTK 26.04 targets gfx906;gfx926;gfx928 by default, compiling each source file three times and increasing build time and memory use. Set export AMDGPU_TARGETS=gfx906 to generate code only for the target architecture.

2.9 gfx906 GPU-aware MPI path

All 12 peer-to-peer checks among the four Z100 cards on Platform B returned hipDeviceCanAccessPeer=1. DTK 26.04 lacks libhsakmt.a, however, so UCX's ROCm transport cannot be built. The alternative is to build OpenMPI 5.0.7 from source with --with-rocm, which provides accelerator: rocm and mpiext: rocm. Use shared memory and disable UCX's network path with --mca pml ob1 --mca btl vader,self and UCX_TLS=sm,self.

2.10 Build differences between the two platforms

Areagfx936/BWgfx906/Z100
Kokkos architectureRequires an explicit gfx936 support patchgfx906 is supported natively
DTK-specific issueNo host-mode regression in 25.04.2DTK 26.04 requires a cuda_wrappers workaround
GPU-aware MPIOpenMPI + UCX + ROCmOpenMPI --with-rocm + shared memory, no UCX ROCm
Build locationBuild completed on the login nodeMoved from the memory-contended login node to a CPU compute node

3. Build configuration and commands

# Use hipcc and enable the Kokkos, HIP, and OpenMP packages
cmake ../lammps-22Jul2025/cmake \
  -D PKG_KOKKOS=ON \
  -D Kokkos_ENABLE_HIP=ON \
  -D PKG_OPENMP=ON \
  -D PKG_KSPACE=ON \
  -D PKG_MANYBODY=ON \
  -D PKG_MOLECULE=ON \
  -D PKG_RIGID=ON \
  -D BUILD_MPI=OFF \
  -D CMAKE_INSTALL_PREFIX=$HOME/softwares/lammps/dcu-install \
  -D CMAKE_CXX_COMPILER=hipcc \
  -D CMAKE_C_COMPILER=gcc \
  -D CMAKE_HIP_COMPILER=hipcc \
  -D HIP_PATH=<DTK_ROOT>/hip

3.1 Single-card build on gfx906/Z100

# Build for gfx906 only; public placeholders are used for shared software and install paths
export AMDGPU_TARGETS=gfx906

cmake -S lammps-22Jul2025/cmake -B build-gfx906 \
  -D PKG_KOKKOS=ON \
  -D Kokkos_ENABLE_HIP=ON \
  -D Kokkos_ARCH_AMD_GFX906=ON \
  -D AMDGPU_TARGETS=gfx906 \
  -D HSA_RUNTIME_LIBRARY=<DTK_ROOT>/hyhal/lib/libhsa-runtime64.so \
  -D AMD_HIP_LIBRARY=<DTK_ROOT>/lib/libamdhip64.so \
  -D PKG_OPENMP=ON -D PKG_KSPACE=ON -D PKG_MANYBODY=ON \
  -D PKG_MOLECULE=ON -D PKG_RIGID=ON \
  -D BUILD_MPI=OFF \
  -D CMAKE_INSTALL_PREFIX=<USER_HOME>/softwares/lammps/dcu-install \
  -D CMAKE_CXX_COMPILER=hipcc -D CMAKE_C_COMPILER=gcc \
  -D CMAKE_HIP_COMPILER=hipcc -D HIP_PATH=<DTK_ROOT>/hip
StepsPlatform B Final Solution
Source codewget https://download.lammps.org/tars/lammps-stable.tar.gz; LAMMPS 22 Jul 2025 Update 5
CMake patchLine 854:-x c++ → -xhip --offload-arch=gfx906
Single card buildBUILD_MPI=OFF, execute
MPI buildOpenMPI 5.0.7 + --with-rocm, LAMMPS changed to BUILD_MPI=ON
Productdcu-install/bin/lmp and dcu-mpi-install/bin/lmp, both about 554 MB

Test job script

#!/bin/bash
#SBATCH -p <partition>
#SBATCH -N 1
#SBATCH --gres=dcu:1
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --time=00:10:00

module load compiler/gcc/9.3.0
module load compiler/dtk/25.04.2

export PATH=$HOME/softwares/lammps/dcu-install/bin:$PATH

lmp -k on g 1 -sf kk -in input.lj

4. Performance results

4.1 Test configuration and workload

CPU test32,000 atoms, 16 OMP threads, 20,000 steps DCU test256,000 atoms, 1 DCU card, 20,000 steps (8 times the number of atoms)

4.2 Benchmark results

Metric CPU (16-core Hygon C86) DCU (1× C-3000 BW) acceleration ratio
Atomic number 32,000 256,000 8×
Simulation steps 20,000 20,000 =
Total time taken 440.81 seconds 21.92 seconds 20.1×
timesteps/s 45.37 912.28 20.1×
Matom-step/s 1.452 233.544 160.8×
Pair calculation time consumption 370.83 seconds (84.1%) 1.69 seconds (7.7%) 219×
Neighbor search takes time 61.83 seconds (14.0%) 3.48 seconds (15.9%) 17.8×
Modify takes time 5.01 seconds (1.1%) 15.34 seconds (70.0%) 0.33×

4.3 DCU resource monitoring

Metric Value
Average usage68%
Peak usage100%
Average power122 W
average temperature54.6 °C

5. Main results

📌 DCU single card processes 8 times the number of atoms, and the total time consumption is only 1/20 of that of CPU
📌 Converted to the same scale, a single DCU card is equivalent to the computing power of about 160-core CPU
📌 Pair compute speedup of 219x, indicating that this compute-intensive workload can significantly benefit from GPU
📌 Modify stage (Verlet points) accounts for 70% on GPU, this is because KOKKOS's full style neighbor list causes large communication overhead
📌 When the DCU is running at full load, the temperature is only 55 °C, the power consumption is 86-122 W, and the energy efficiency ratio is extremely high

5.1 gfx906/Z100 Correctness and Energy Conservation

platform B uses the same input to run Kokkos/HIP and CPU/OpenMP 16 threads respectively to avoid judging the success of the transplant based on performance alone.

Calculation exampleMetricGPU (Kokkos)CPU (OpenMP)Deviation
LJ 10000 stepsTotEng-4.62126-4.621190.0015%
LJ 10000 stepsTemp0.697230.695910.19%
EAM 1000 stepsTotEng-106640.19-106640.19Printing accuracy is consistent within
EAM 1000 stepsTemp796.34045796.34045Printing accuracy is consistent within
EAM 1000 stepsE_pair-109934.01-109934.01Printing accuracy is consistent within

LJ comes from the parallel floating point summation order. In the NVE convergence test, LJ can always change from -4.6203 to -4.6213 after equilibrium, with a drift less than 0.02%; EAM changes from -106640.66 to -106640.19, with a drift less than 0.0004%.

5.2 gfx906/Z100 calculation example coverage

Calculation exampleBag/StyleGPUCPU
in.ljpair lj/cutPassedPassed
in.eampair eam / MANYBODYPassedPassed
in.chainbond fene / MOLECULEPassedNot tested
in.rhodopppm/kk + CHARMM + SHAKE / KSPACErequires neigh halfPassed

5.3 gfx906/Z100 single card weak expansion

Calculation exampleAtomic numberstep/sMatom-step/sObservation
LJ32K223471.5—
LJ256K472120.9—
LJ2M66.9137.1Enters the saturation zone
LJ16.4M8.18134.0Stable saturation
EAM32K62520.0—
EAM256K180.646.2—
EAM2M25.251.5approaching saturation

5.4 gfx906/Z100 multi-card strong expansion (2M atoms)

Calculation exampleGPU numbernon-GPU-awareGPU-awareaware improverelative to 1 card
LJ167.367.7—1.00×
LJ244.1(0.65×)91.42.07×1.35×
LJ478.1125.81.61×1.86×
EAM125.525.5—1.00×
EAM222.5(0.88×)39.91.77×1.56×
EAM440.559.51.47×2.33×

GPU-aware MPI changes two-card performance from a slowdown to a 1.35–1.56× speedup, and four-card performance reaches 1.86–2.33×. This result is in stark contrast to Platform A's multi-card non-breakout single card, indicating that the communication stack, P2P topology, problem size, and site configuration are all part of the conclusion.

5.5 Cross-platform performance comparison

Metricgfx906/Z100gfx936/BWObservation
Architecture / CU / MemoryVega 20 / 64 / 16 GBgfx936 / 80 / 64 GBDifferent hardware and capacity
LJ 256K472 step/sabout 912 steps/sAbout 52%
LJ saturation throughputAbout 134 Matom-step/sAbout 80 Matom-step/sPlatform B is about 1.68× faster
EAM 2M25.2 step/sAbout 21 step/sPlatform B is about 1.2× faster
Two-card LJ scaling1.35×0.71–0.79×Platform B is better

The two environments do not have the same input, version and communication stack. This table preserves only the trends reported in the original results and cannot be used as a strict hardware ranking. Platform B's same-node PCIe P2P efficiency is an important factor in multi-card performance.

6. GPU package and Kokkos performance comparison

6.1 Test purpose

This section evaluates whether the GPU package offers a more efficient acceleration path than Kokkos. The GPU package uses hand-written optimized HIP kernels that may theoretically be more efficient than Kokkos' general abstraction.

6.2 GPU package compilation repair

Key repair points
  • AMDGPU_TARGETS="gfx936" and GPU_TARGETS="gfx936": Prevent DTK from automatically adding multiple architectures resulting in --offload-arch leaked to g++
  • HIP_USE_DEVICE_SORT=OFF: Avoid linking hip::device, prevent -xhip flag is passed to g++

6.3 Performance comparison results (32K atoms, 1000 steps)

Metric GPU Package + HIP Kokkos + HIP Kokkos Advantages
Total time taken 3.47 seconds 0.55 seconds 6.3×
timesteps/s 288.1 1804.1 6.3×
Matom-step/s 37.8 236.5 6.3×

6.4 Performance comparison results (256K atoms, 1000 steps)

Metric GPU Package + HIP Kokkos + HIP Kokkos Advantages
Total time taken 14.73 seconds 3.20 seconds 4.6×
timesteps/s 67.9 312.5 4.6×
Matom-step/s 71.2 327.7 4.6×

6.5 Why the GPU package is slower

📌 Neighbor list built on the CPU: The the GPU package reports “Neigh mode: Hybrid (binning on host).” Each rebuild transfers data from CPU to GPU, while Kokkos builds the neighbor list entirely on the GPU
📌 GPU package design favors multi-GPU: The single-GPU scenario is expensive and requires an additional fix gpu command to manage resources
📌 Kokkos’ GPU execution scope is more complete: pair, neighbor, and communication work all run on the GPU, reducing data transmission between the host and the device

6.6 Hygon native compiler confirmation

Inspection of the DTK toolchain confirms that this build uses the Hygon native compiler:

# hipcc ultimately invokes Hygon's native dcc compiler
hipcc → <DTK_ROOT>/dcc/bin/dcc
           ↓
  Hygon ISA bitcode: oclc_isa_version_936.bc

Kokkos' AMD_GFX936 is an architecture label. Compilation uses Hygon's dcc compiler and Hygon ISA bitcode, so the resulting code is compiled natively for the DCU.

7. Hotspot analysis and code optimization

7.1 HIP Hotspot Tracking

Hotspots in the 1,048,576-atom dynamic LJ benchmark were analyzed with DTK 25.04.2 hipprof --hip-trace --stats. of the 1,048,576-atom dynamic LJ benchmark.

KernelNumber of callsTotal timeproportion
LJ pair force5,19412.877 s64.92%
full neighbor-list build2635.523 s27.85%
fused NVE integrate5,1941.091 s5.50%
all remaining kernels-0.344 s1.74%

Pair and Neighbor account for 92.8% of device time and are the only worthwhile optimization targets.

7.2 PMC hardware counter analysis

PMC counters were collected for the Pair and Neighbor hotspots:

MetricLJ pair forceNeighbor build
VALU command70,299,737433,198,843
VMEM read command9,440,57419,685,740
VMEM write command65,5369,601,271
LDS command079,223,620
LDS wait command014,343,770
L2 hit rate95.10%83.77%
TCP data-stall cycle3,537,466467,603,303

Pair does not use LDS and has a high L2 hit rate; Neighbor has a large number of LDS and TCP stall activities, which is the focus of the next optimization step.

7.3 Pair workgroup scan (negative optimization)

tests workgroup 64/128/256/512/1024, staggered operation every three rounds:

Workgroupmean (step/s)relative to 1024
64228.706-8.58%
128230.307-7.94%
256233.920-6.50%
512241.857-3.32%
1024250.173Baseline

improves monotonically as the workgroup increases, and the value of 1024 automatically selected by Kokkos is already the best value in the test. Experimental code has been rolled back.

7.4 Neighbor transpose optimization (+13.6%)

Runtime options neigh/transpose on, no need to recompile, three rounds of interleaved verification:

ModeRun 1Run 2Run 3meanImprove
transpose on284.280284.007284.166284.151+13.61%
transpose off249.724250.186250.375250.095Baseline

changes the memory layout of the neighbor list on the GPU (row major → column major) and optimizes the cache access mode of the pair kernel. Attention transpose on is related to the system size: +13.6% for 1M atoms, but slightly slower -1.2% for 32K atoms.

7.5 Neighbor BPT optimization (+4.31%)

tests the number of bins per team in the neighbor build (BPT=1/2/4/8), three rounds of staggered runs:

BPTmean (step/s)relative to the default
1260.742+4.31%
2 (default)249.974Baseline
4258.720+3.50%
8260.992+4.41%

The default BPT=2 performs worst among the four configurations. Setting BPT=1 improves performance by 4.31% with a one-line change (const int factor = 1;); this setting is retained in the source.

8. CPU and DCU performance comparison

8.1 Limitations of the original “160 cores” comparison

The original report equates one DCU card with a 160-core CPU, based on results from different system sizes:

PlatformAtomic numbertimesteps/sMatom-step/s
CPU (16 OMP threads)32,00045.371.452
DCU (original)256,000912.28233.544

The normalized throughput ratio is 233.544 / 1.452 = 160.8×, but the CPU and DCU results use different atom counts. The larger DCU case achieves higher GPU utilization, so this cross-scale ratio overstates the speedup for a like-for-like comparison.

8.2 Fair comparison of same scale (32,000 atoms)

Platformstep/sMatom-step/sacceleration ratio
CPU (16 OMP threads)52.921.6931×
DCU (1× C-3000 BW)3,206.09102.59560.6×

At the same scale, one DCU card is about 60× faster, rather than 160×. Follow-up tests measured GPU utilization of 20–41%, depending on the pair style; CPU-side framework overhead was the main bottleneck.

8.3 How problem size changes the GPU speedup

📌 GPU speedup grows with system size: 60× at 32K atoms and an estimated ~172× at 1M, as fixed overhead is amortized, GPU occupancy rises, and the compute-to-memory-access ratio improves
📌 GPU utilization is only 20% on the LJ benchmark: Pair calculations account for 6.8% of runtime, while Modify (CPU-side framework work) accounts for 76%. LJ is too lightweight to be a representative GPU optimization target
📌 GPU utilization reaches 41% on the EAM benchmark: Pair calculations account for 36.8% of runtime, and EAM performs 20.7× as much computation as LJ. EAM is a more representative GPU benchmark

9. Optimization experiments

9.1 Effective optimization

OptimizationBenchmarkProfitPrinciplePrice
neigh/transpose on LJ 1M +13.61% Neighbor list row-major → column-major, optimized pair kernel cache access mode Runtime option, zero cost
neigh/transpose on EAM 1M +3.18% Same as above, but EAM calculation is more intensive and memory access optimization has diminishing returns Same as above
Neighbor BPT=1 LJ 1M +4.31% Each team processes 1 bin (original default 2), halving LDS usage → higher occupancy One line of source code modification
Neighbor BPT=1 EAM 1M +0.60% Neighbor only accounts for 4.6% of the total EAM time, and there is very little room for optimization Same as above

9.2 Invalid optimization (negative optimization or zero benefit)

OptimizationProfitCause AnalysisProcessing
Pair workgroup 64 -8.58% The smaller the workgroup, the fewer the number of waves per CU, and the degree of parallelism decreases code has been rolled back
Pair workgroup 128 -7.94% performance increases monotonically with workgroup, Kokkos automatically selects 1024 Correct code has been rolled back
Pair workgroup 256 -6.50%
Pair workgroup 512 -3.32%
DTK 25.04.4 upgrade -0.23% No performance difference between compiler/runtime version iterations does not upgrade
DTK 26.04 upgrade -0.11%
fix nve/kk mask fast path -0.12% Branch prediction fails, additional condition judgment increases overhead code has been rolled back
fix nve/kk single type quality -0.49% Type checking overhead exceeds computational simplification benefits code has been rolled back
nbin/atoms/per/bin scan Unable to test requires neigh/thread on, which will change the pair algorithm and confuse comparison Give up
2 GPU (no GPU-aware) 0.42× Comm accounts for 42%, data is copied back and forth between CPU and GPU Upgrade to GPU-aware
2 GPU (GPU-aware, LJ) 0.71-0.79× Comm down to 8-30%, but LJ calculation is too light to dilute communication Use EAM
2 GPU (GPU-aware, EAM 4M) 0.78× EAM is more computationally intensive, the 2/1 ratio has increased from 0.78 to 0.87, the trend is improving but has not exceeded 1.0 Need larger scale
2 GPU (GPU-aware, EAM 8M) 0.87×

9.3 Summary of key experiences

📌 LJ benchmark is not suitable for GPU optimization: GPU is only 20% active, Pair accounts for 6.8%. The optimization effect is amplified (transpose +13.6% on EAM only +3.2%)
📌 EAM is a more reasonable GPU benchmark: GPU active 41%, Pair accounts for 36.8%, calculation amount 20.7× LJ
📌 Memory-access optimizations yield less as compute intensity increases: transpose +13.6% on LJ, only +3.2% on EAM
📌 GPU utilization is limited by Modify overhead: Modify consumes 57–76% of runtime as CPU-side framework overhead, which is difficult to reduce
📌 Multi-GPU requires GPU-aware MPI + compute intensive potential function + large number of atoms: Three conditions are indispensable

10. MPI build and GPU-aware communication

10.1 Background

LAMMPS multi-GPU requires MPI support. There is an ABI mismatch between the OpenMPI that comes with the system (compiled with GCC 8.5.0) and the current GCC 9.3.0, resulting in MPI_Finalize when double free crash. In addition, non-GPU-aware MPI requires CPU↔GPU data copy during each communication step, which results in huge communication overhead.

10.2 Compilation environment

ComponentVersionSource
GCC9.3.0module load compiler/gcc/9.3.0
DTK25.04.2module load compiler/dtk/25.04.2
UCX1.18.0module load mpi/ucx/1.18.0/dtk-25.04/shca
OpenMPI5.0.7source code compilation

10.3 Compilation steps

# Download and build OpenMPI 5.0.7
wget https://download.open-mpi.org/release/open-mpi/v5.0/openmpi-5.0.7.tar.gz
tar xzf openmpi-5.0.7.tar.gz && cd openmpi-5.0.7

# Load dependencies
module purge
module load compiler/gcc/9.3.0
module load compiler/dtk/25.04.2
module load <UCX_MODULE>

# Configure UCX and ROCm GPU support
./configure --prefix=$HOME/softwares/mpi/install \
  --enable-mpi-fortran=no \
  --with-ucx=<UCX_ROOT> \
  --with-rocm=<DTK_ROOT>

make -j32 && make install

10.4 GPU-aware MPI components

ComponentDescriptionVerification command
accelerator: rocmGPU direct memory access, avoiding CPU↔GPU copyompi_info --all | grep rocm
pml: ucxUCX point-to-point communication layer, supports InfiniBand/RoCEompi_info --all | grep 'pml: ucx'
osc: ucxUCX Unilateral communication (RDMA)ompi_info --all | grep 'osc: ucx'
mpiext: rocmROCm MPI extensionompi_info --all | grep mpiext

10.5 LAMMPS MPI build

# Build LAMMPS with the locally built MPI
cmake -S lammps-22Jul2025/cmake -B build-mpi \
  -C cmake/presets/hygon-dcu-gfx936.cmake \
  -D CMAKE_CXX_COMPILER=$DTKROOT/bin/hipcc \
  -D BUILD_MPI=ON \
  -D MPI_CXX_COMPILER=$HOME/softwares/mpi/install/bin/mpicxx \
  -D MPI_CXX_INCLUDE_PATH=$HOME/softwares/mpi/install/include \
  -D MPI_CXX_LIBRARIES=$HOME/softwares/mpi/install/lib/libmpi.so

10.6 Runtime configuration

# Key Slurm job settings
module load <UCX_MODULE>
export PATH=$HOME/softwares/mpi/install/bin:$PATH
export LD_LIBRARY_PATH=$HOME/softwares/mpi/install/lib:$LD_LIBRARY_PATH

# Enable GPU-aware communication with -pk kokkos gpu/aware on
mpirun -np 2 lmp -in input -k on g 1 -sf kk \
  -pk kokkos neigh/transpose on gpu/aware on

10.7 MPI version comparison

MPI version2 GPU step/sCommshutdownGPU-aware
System MPI (GCC 8.5, no UCX)104.6542%crashNo
Self-compiled MPI (GCC 9.3, no UCX)104.6542%✅No
GPU-aware MPI (GCC 9.3 + UCX+ROCm)176.9540%✅Yes

GPU-aware MPI improves two-card performance by 69% (105 → 177 steps/s) and reduces communication time from 20.2 s to 11.4 s (−43%). The UCX + ROCm path avoids copying data through the CPU.

10.8 UCX selection guidance

DTK 25.04.2 does not include UCX. The tested <UCX_MODULE> provides UCX components with ROCm transport support. Building OpenMPI 5.0.7 with --with-ucx and --with-rocm links these components and enables GPU-aware communication.

UCX module path:<UCX_ROOT>

ROCm path:<DTK_ROOT>

10.9 gfx906/Z100: Same-node solution without UCX ROCm

Platform B lacks the static library libhsakmt.a, so the UCX ROCm transport cannot be built. OpenMPI can still recognize device pointers through its ROCm accelerator extension and use shared memory for communication among cards on the same node:

./configure --prefix=<USER_HOME>/softwares/mpi/ompi-install \
  --enable-mpi-fortran=no \
  --with-rocm=<DTK_ROOT> \
  --with-slurm

export UCX_TLS=sm,self
mpirun --mca pml ob1 --mca btl vader,self \
  -np <GPU_COUNT> lmp -in input \
  -k on g 1 -sf kk -pk kokkos gpu/aware on

This build exposes accelerator: rocm and mpiext: rocm. It was used for intra-node peer-to-peer tests with one to four cards; it does not validate cross-node RDMA.

11. Artifact directory structure

~/softwares/lammps/
├── lammps-22Jul2025/ # Source code (Git repository)
│ ├── src/KOKKOS/
│ │ ├── npair_kokkos.cpp # Neighbor BPT=1 Optimization reserved
│ │ └── pair_kokkos.h # Pair workgroup experiment has been rolled back
│ ├── lib/kokkos/
│ │ ├── cmake/ # gfx926/928/936/938 architecture support
│ │ └── core/src/HIP/ # wavefront 64 + system assigned properties
│ ├── cmake/presets/
│ │ └── hygon-dcu-gfx936.cmake # Reproducible construction preset
│ └── tools/hygon/ # Build scripts, benchmarks, analysis tools
│ ├── build-gfx936.sh # Version isolation build
│ ├── analyze-pmc.py # PMC automatic reduction
│ ├── benchmark/ # Benchmark and Slurm jobs
│ └── results/ # Experimental conclusion Markdown
├── install-dtk-25.04.2-gfx936/ # Baseline binary
├── install-pair-wg-*-gfx936/ # Pair experimental binary (reserved)
├── install-neighbor-bpt-*-gfx936/ # Neighbor experimental binary (reserved)
└── dcu-test/ # Original profiling data
├── hipprof-<run-id>/ # HIP trace data
├── pmc-<run-id>/ # PMC Original CSV
├── pair-workgroup-<run-id>/ # Pair scan log
└── neighbor-bpt-<run-id>/ # Neighbor scan log

12. Conclusions and scope

Recommended solution

Within the range of hardware, software versions and workloads described in this report,KOKKOS + HIP + neigh/transpose on achieved the best overall performance.

12.1 Architecture and compilation

✅ gfx936 native compilation: hipcc → dcc (Hygon DCU C compiler), generates Hygon ISA code
✅ Kokkos architecture supports: gfx926/928/936/938 registered to cmake and HIP core file
✅ GPU-aware MPI is available:OpenMPI 5.0.7 + UCX 1.18.0 DTK + ROCm

12.2 Effective optimization

✅ neigh/transpose on: LJ +13.6%, EAM +3.2%; it is a runtime option, no need to modify the source code
✅ KOKKOS outperforms GPU package: EAM 1M atoms of the same scale, 72.9 vs 39.6 step/s (1.84×)

12.3 Invalid optimization (rolled back)

❌ Pair workgroup 64-512: slow 3.3-8.6%, Kokkos automatically selects 1024 correct
❌ DTK upgrade: 25.04.4/26.04 No performance difference
❌ fix nve/kk source code: mask fast path -0.12%, single type quality -0.49%
❌ BPT=1: LJ +4.3%, but EAM only +0.6%; rolled back, current evidence is insufficient to support this as a stable and generalizable optimization

12.4 Multi-GPU scaling

Pair StyleMode1 GPU2 GPU4 GPU2/14/1
EAM 1MKOKKOS76.557.145.20.750.59
EAM 4MKOKKOS21.216.5—0.78—
EAM 8MKOKKOS10.28.9—0.87—
EAM 16MKOKKOS5.24.7—0.90—
LJ 1MKOKKOS277.7198.3—0.71—
Tersoff 1.1MKOKKOS195.0123.2—0.63—
EAM 1MGPU package39.6————

EAM 2/1 ratio 0.78→0.87→0.90, trend improving but not breaking 1.0.GPU package has good scalability (2/1=1.32) but single card is slow (~1.8× KOKKOS).

12.5 Fair comparison of CPU and DCU performance

vs.Original reportAfter correctionQuestion
acceleration ratio160×60×Original cross-scale comparison (32K CPU vs 256K DCU) unfair
Acceleration trend—As the number of atoms increases32K: 60× → 1M: ~172× (estimate)

12.6 Single-GPU scaling

Atomic numberstep/sMatom-step/sPair%memory
1M72.576.036.7%0.16GB
4M21.285.032.8%0.64GB
8M10.281.933.8%1.2GB
16M5.282.632.2%2.4GB
25M3.280.232.0%3.8GB
32M———OOM

Throughput is stable at 80–85 Matom-step/s with near-linear scaling. The 32M-atom case runs out of memory because the neighbor list initially allocates 2,000 entries per atom, exceeding 64 GB. Memory capacity becomes the limit before compute throughput.

12.7 Key lessons

⚠️ LJ is not suitable for GPU benchmark: The GPU is only 20% active, and the optimization effect is amplified
⚠️ EAM is a reasonable GPU benchmark: GPU active 41%, Pair accounts for 36.8%
⚠️ Device utilization during the active computing phase can approach 100%(HIP trace sampling), but end-to-end performance is still limited by CPU-side framework overhead (Modify accounts for 57-76%)
⚠️ Memory capacity limits scale before compute throughput: the 32M-atom neighbor-list allocation requires 256 GB, exceeding the 64 GB available
⚠️ Three requirements for multi-GPU scaling: GPU-aware MPI, a compute-intensive potential, and a large system (>500K atoms per GPU) are all necessary.
⚠️ GPU package vs KOKKOS: The GPU package has good scalability but a single card is slow. KOKKOS is fast with a single card but poor with multiple GPUs.

12.8 gfx906/Z100 limitations and workarounds

CHARMM/Kokkos neighbor list

dihedral charmm/kk implements only HALF neighbor lists, while GPU mode defaults to neighflag=FULL. atom_style full together with lj/charmm/coul/long also forces FULL lists, creating a conflict. Running the post-rhodo case on the GPU with -pk kokkos neigh half matches the CPU result. The missing FULL-list implementation remains an upstream limitation.

CategoryRestrictionsProcess/Affect
DTK 26.04host mode injection cuda_wrappersChange lmp target to -xhip --offload-arch=gfx906
DTK 26.04Missing libhsakmt.aUnable to build UCX ROCm; use OpenMPI --with-rocm
HSA runtimesymbolic chain depends on /opt/hyhalExplicitly specify the real library in the shared software area
Compilationgenerates gfx906/926/928 three-architecture by defaultFixed AMDGPU_TARGETS=gfx906
Hardwaregfx906 No FP8/BF16 support or MFMA matrix coresThis does not affect the molecular-dynamics tests in this report; results should not be extrapolated to matrix workloads
GPU memory16 GB per card; LJ runs are limited to about 16 million atomsLarger systems require multiple cards or lower memory use
SchedulingThe test partition rejects requests for more than 8 CPU coresThe multi-card test used 8 CPU cores in total; the partition name is redacted

12.9 Lessons from the two platforms

✅ First distinguish the architecture support status: gfx906 has been natively supported by Kokkos; gfx936 needs to register the architecture, wavefront and system allocation attributes, and the two cannot reuse the same patch assumption.
⚠️ DTK upgrades may have no performance gain or may introduce build regressions: Performance on platform A was nearly unchanged between DTK 25.04.4 and 26.04, but 26.04 on platform B introduces the cuda_wrappers host-mode problem.
✅ Compile HIP workloads on compute nodes: Memory use on a shared login node can fluctuate substantially with concurrent users. Submit resource requests within the site CPU-to-memory limit, and propagate the job exit code explicitly.
✅ GPU-aware does not have to rely heavily on UCX ROCm: Platform B provides same-node GPU-aware communication through OpenMPI --with-rocm and shared memory; Platform A uses the UCX+ROCm path.
⚠️ Build OpenMPI locally to control the transport layer: Platform B uses pml ob1, btl vader,self, and UCX_TLS=sm,self to avoid GID assertions in unsupported InfiniBand/UCX paths.
⚠️ Multi-card scaling requires GPU-aware MPI, roughly 500K or more atoms per card, and a sufficiently compute-intensive potential. Memory capacity may become the limit before compute throughput.

13. Upstream improvement proposals

13.1 Kokkos upstream

ContributionContentfileStatus
gfx926/928/936/938 architecture supportAdds Hygon DCU architecture to Kokkos cmake and HIP corekokkos_arch.cmake, KokkosCore_config.h.in, Kokkos_HIP_Instance.hpp, Kokkos_HIP_IsXnack.hppprototype has been verified, waiting for upstream review
wavefront 64 supportgfx936 uses 64-wide wavefront, same as gfx906Kokkos_HIP_Instance.hppLocal implementation verified
System allocated access attributeHygon DCU does not support direct access to system memoryKokkos_HIP_IsXnack.hppLocal implementation verified

13.2 LAMMPS upstream

ContributeContentStatus
build presetcmake/presets/hygon-dcu-gfx936.cmake: Reproducible Hygon DCU build configurationTo be submitted
Performance BenchmarkKOKKOS vs GPU package comparison data, multi-GPU scaling analysisavailable for document reference
Recommended configurationneigh/transpose on has +3~14% improvement over gfx936Runtime options, no code required

13.3 ROCm/HIP compatibility ecosystem

ContributionContentStatus
gfx936 performance dataSingle/multi-card performance of LJ, EAM, Tersoff, ReaxFF on 80 CUavailable for document reference
GPU-aware MPI solutionOpenMPI 5.0.7 + UCX 1.18.0 DTK + ROCm compilation methodDocumented
Verification methodPMC analysis tools, HIP trace workflow, dynamic benchmarktools/hygon/