This report brings together two independent validations built around the same engineering workflow. Platform A covers gfx936 architecture support, GPU package comparisons, and source optimizations; Platform B covers native gfx906/Z100 builds, correctness, and multi-card scaling. Locations, cluster aliases, SSH addresses, partition and node names, and absolute user paths are redacted.
--with-rocm)
SchedulerSlurm 3.2.10; CPU build and 1–4 DCU test partition names redacted
Configuration: -D PKG_GPU=ON -D GPU_API=HIP -D HIP_ARCH=gfx936
First result:Compiled successfully, but hung when running
Initial diagnosis:DTK's CMake configuration applies the -xhip flag to the GPU package, causing hipcc to compile every .cpp file for the device. The resulting executable hung during HIP initialization.
After the fix:Setting AMDGPU_TARGETS="gfx936" and disabling HIP_USE_DEVICE_SORT allowed the build to compile and run. Performance tests nevertheless showed the GPU package to be 4.6–6.3× slower than Kokkos (see the comparison in Section 6).
Configuration: -D PKG_KOKKOS=ON -D Kokkos_ENABLE_HIP=ON
Result:The build succeeds, but execution fails with running kernels compiled for gfx906 on gfx936.
Analysis:Kokkos 4.6.2 does not recognize gfx936 and falls back to gfx906. Because the gfx906 and gfx936 instruction sets are incompatible, execution fails with hipErrorInvalidKernelFile.
Approach:Add gfx936 support explicitly to Kokkos' CMake configuration and HIP backend.
AMD_GFX936 and its gfx936 target name. List gfx936 before gfx906 so automatic detection selects the correct architecture first, and register the numeric alias 936.#cmakedefine KOKKOS_ARCH_AMD_GFX936.defined(KOKKOS_ARCH_AMD_GFX936) in the HIPTraits condition that sets WarpSize. Hygon BW DCUs use a wavefront size of 64, as do gfx906 and gfx908.defined(KOKKOS_ARCH_AMD_GFX936) to the gpu_arch_can_access_system_allocations branch that returns false. Hygon DCUs do not support access to system allocations.Unlike gfx936, gfx906 (Vega 20) has been natively supported by Kokkos 4.6.2, so there is no need to modify the four architecture files of Kokkos and directly enable Kokkos_ARCH_AMD_GFX906=ON and AMDGPU_TARGETS=gfx906. This reduces architecture porting costs, but DTK 26.04 and the cluster toolchain expose another set of build issues.
At about 98% of the build, main.cpp failed because cuda_wrappers/algorithm reported unknown type name '__host__'. In LAMMPS Update 5, cmake/CMakeLists.txt:854 forces the lmp target to -x c++ to address an nlohmann/json HIP SFINAE issue. DTK 26.04's hipcc still injects CUDA wrappers in host mode, causing the standard <algorithm> header to be parsed incorrectly.
-x c++ option with -xhip --offload-arch=gfx906:target_compile_options(lmp PRIVATE "$<$<COMPILE_LANGUAGE:CXX>:SHELL:-xhip --offload-arch=gfx906>")The same issue is not reproduced on Platform A in DTK 25.04.2 and is therefore logged as a DTK 26.04 regression rather than a general LAMMPS defect.
The libhsa-runtime64.so symlink resolves through the DCU driver directory /opt/hyhal/, which is absent from the CPU build node. CMake therefore cannot find the HSA runtime. Set HSA_RUNTIME_LIBRARY explicitly to the actual shared-library path; the public commands use the placeholder <DTK_ROOT>/hyhal/lib/libhsa-runtime64.so.
On Platform B, the system /usr/bin/make references a missing ELF interpreter at /opt/glibc-2.25/lib/ld-linux.so.2; load compiler/make/4.4 instead. The login node has about 251 GB of shared memory for roughly 249 users, with only about 46 GB available during testing. When hipcc compiled roughly 1,000 C++ files, dcc was killed by the OOM killer. The build was moved to the CPU compute queue and submitted with 32 cores and about 111 GB of exclusive memory.
DTK 26.04 targets gfx906;gfx926;gfx928 by default, compiling each source file three times and increasing build time and memory use. Set export AMDGPU_TARGETS=gfx906 to generate code only for the target architecture.
All 12 peer-to-peer checks among the four Z100 cards on Platform B returned hipDeviceCanAccessPeer=1. DTK 26.04 lacks libhsakmt.a, however, so UCX's ROCm transport cannot be built. The alternative is to build OpenMPI 5.0.7 from source with --with-rocm, which provides accelerator: rocm and mpiext: rocm. Use shared memory and disable UCX's network path with --mca pml ob1 --mca btl vader,self and UCX_TLS=sm,self.
| Area | gfx936/BW | gfx906/Z100 |
|---|---|---|
| Kokkos architecture | Requires an explicit gfx936 support patch | gfx906 is supported natively |
| DTK-specific issue | No host-mode regression in 25.04.2 | DTK 26.04 requires a cuda_wrappers workaround |
| GPU-aware MPI | OpenMPI + UCX + ROCm | OpenMPI --with-rocm + shared memory, no UCX ROCm |
| Build location | Build completed on the login node | Moved from the memory-contended login node to a CPU compute node |
# Use hipcc and enable the Kokkos, HIP, and OpenMP packages cmake ../lammps-22Jul2025/cmake \ -D PKG_KOKKOS=ON \ -D Kokkos_ENABLE_HIP=ON \ -D PKG_OPENMP=ON \ -D PKG_KSPACE=ON \ -D PKG_MANYBODY=ON \ -D PKG_MOLECULE=ON \ -D PKG_RIGID=ON \ -D BUILD_MPI=OFF \ -D CMAKE_INSTALL_PREFIX=$HOME/softwares/lammps/dcu-install \ -D CMAKE_CXX_COMPILER=hipcc \ -D CMAKE_C_COMPILER=gcc \ -D CMAKE_HIP_COMPILER=hipcc \ -D HIP_PATH=<DTK_ROOT>/hip
# Build for gfx906 only; public placeholders are used for shared software and install paths
export AMDGPU_TARGETS=gfx906
cmake -S lammps-22Jul2025/cmake -B build-gfx906 \
-D PKG_KOKKOS=ON \
-D Kokkos_ENABLE_HIP=ON \
-D Kokkos_ARCH_AMD_GFX906=ON \
-D AMDGPU_TARGETS=gfx906 \
-D HSA_RUNTIME_LIBRARY=<DTK_ROOT>/hyhal/lib/libhsa-runtime64.so \
-D AMD_HIP_LIBRARY=<DTK_ROOT>/lib/libamdhip64.so \
-D PKG_OPENMP=ON -D PKG_KSPACE=ON -D PKG_MANYBODY=ON \
-D PKG_MOLECULE=ON -D PKG_RIGID=ON \
-D BUILD_MPI=OFF \
-D CMAKE_INSTALL_PREFIX=<USER_HOME>/softwares/lammps/dcu-install \
-D CMAKE_CXX_COMPILER=hipcc -D CMAKE_C_COMPILER=gcc \
-D CMAKE_HIP_COMPILER=hipcc -D HIP_PATH=<DTK_ROOT>/hip
| Steps | Platform B Final Solution |
|---|---|
| Source code | wget https://download.lammps.org/tars/lammps-stable.tar.gz; LAMMPS 22 Jul 2025 Update 5 |
| CMake patch | Line 854:-x c++ → -xhip --offload-arch=gfx906 |
| Single card build | BUILD_MPI=OFF, execute |
| MPI build | OpenMPI 5.0.7 + --with-rocm, LAMMPS changed to BUILD_MPI=ON |
| Product | dcu-install/bin/lmp and dcu-mpi-install/bin/lmp, both about 554 MB |
#!/bin/bash #SBATCH -p <partition> #SBATCH -N 1 #SBATCH --gres=dcu:1 #SBATCH --ntasks-per-node=1 #SBATCH --cpus-per-task=16 #SBATCH --time=00:10:00 module load compiler/gcc/9.3.0 module load compiler/dtk/25.04.2 export PATH=$HOME/softwares/lammps/dcu-install/bin:$PATH lmp -k on g 1 -sf kk -in input.lj
| Metric | CPU (16-core Hygon C86) | DCU (1× C-3000 BW) | acceleration ratio |
|---|---|---|---|
| Atomic number | 32,000 | 256,000 | 8× |
| Simulation steps | 20,000 | 20,000 | = |
| Total time taken | 440.81 seconds | 21.92 seconds | 20.1× |
| timesteps/s | 45.37 | 912.28 | 20.1× |
| Matom-step/s | 1.452 | 233.544 | 160.8× |
| Pair calculation time consumption | 370.83 seconds (84.1%) | 1.69 seconds (7.7%) | 219× |
| Neighbor search takes time | 61.83 seconds (14.0%) | 3.48 seconds (15.9%) | 17.8× |
| Modify takes time | 5.01 seconds (1.1%) | 15.34 seconds (70.0%) | 0.33× |
| Metric | Value |
|---|---|
| Average usage | 68% |
| Peak usage | 100% |
| Average power | 122 W |
| average temperature | 54.6 °C |
platform B uses the same input to run Kokkos/HIP and CPU/OpenMP 16 threads respectively to avoid judging the success of the transplant based on performance alone.
| Calculation example | Metric | GPU (Kokkos) | CPU (OpenMP) | Deviation |
|---|---|---|---|---|
| LJ 10000 steps | TotEng | -4.62126 | -4.62119 | 0.0015% |
| LJ 10000 steps | Temp | 0.69723 | 0.69591 | 0.19% |
| EAM 1000 steps | TotEng | -106640.19 | -106640.19 | Printing accuracy is consistent within |
| EAM 1000 steps | Temp | 796.34045 | 796.34045 | Printing accuracy is consistent within |
| EAM 1000 steps | E_pair | -109934.01 | -109934.01 | Printing accuracy is consistent within |
LJ comes from the parallel floating point summation order. In the NVE convergence test, LJ can always change from -4.6203 to -4.6213 after equilibrium, with a drift less than 0.02%; EAM changes from -106640.66 to -106640.19, with a drift less than 0.0004%.
| Calculation example | Bag/Style | GPU | CPU |
|---|---|---|---|
in.lj | pair lj/cut | Passed | Passed |
in.eam | pair eam / MANYBODY | Passed | Passed |
in.chain | bond fene / MOLECULE | Passed | Not tested |
in.rhodo | pppm/kk + CHARMM + SHAKE / KSPACE | requires neigh half | Passed |
| Calculation example | Atomic number | step/s | Matom-step/s | Observation |
|---|---|---|---|---|
| LJ | 32K | 2234 | 71.5 | — |
| LJ | 256K | 472 | 120.9 | — |
| LJ | 2M | 66.9 | 137.1 | Enters the saturation zone |
| LJ | 16.4M | 8.18 | 134.0 | Stable saturation |
| EAM | 32K | 625 | 20.0 | — |
| EAM | 256K | 180.6 | 46.2 | — |
| EAM | 2M | 25.2 | 51.5 | approaching saturation |
| Calculation example | GPU number | non-GPU-aware | GPU-aware | aware improve | relative to 1 card |
|---|---|---|---|---|---|
| LJ | 1 | 67.3 | 67.7 | — | 1.00× |
| LJ | 2 | 44.1(0.65×) | 91.4 | 2.07× | 1.35× |
| LJ | 4 | 78.1 | 125.8 | 1.61× | 1.86× |
| EAM | 1 | 25.5 | 25.5 | — | 1.00× |
| EAM | 2 | 22.5(0.88×) | 39.9 | 1.77× | 1.56× |
| EAM | 4 | 40.5 | 59.5 | 1.47× | 2.33× |
GPU-aware MPI changes two-card performance from a slowdown to a 1.35–1.56× speedup, and four-card performance reaches 1.86–2.33×. This result is in stark contrast to Platform A's multi-card non-breakout single card, indicating that the communication stack, P2P topology, problem size, and site configuration are all part of the conclusion.
| Metric | gfx906/Z100 | gfx936/BW | Observation |
|---|---|---|---|
| Architecture / CU / Memory | Vega 20 / 64 / 16 GB | gfx936 / 80 / 64 GB | Different hardware and capacity |
| LJ 256K | 472 step/s | about 912 steps/s | About 52% |
| LJ saturation throughput | About 134 Matom-step/s | About 80 Matom-step/s | Platform B is about 1.68× faster |
| EAM 2M | 25.2 step/s | About 21 step/s | Platform B is about 1.2× faster |
| Two-card LJ scaling | 1.35× | 0.71–0.79× | Platform B is better |
The two environments do not have the same input, version and communication stack. This table preserves only the trends reported in the original results and cannot be used as a strict hardware ranking. Platform B's same-node PCIe P2P efficiency is an important factor in multi-card performance.
This section evaluates whether the GPU package offers a more efficient acceleration path than Kokkos. The GPU package uses hand-written optimized HIP kernels that may theoretically be more efficient than Kokkos' general abstraction.
AMDGPU_TARGETS="gfx936" and GPU_TARGETS="gfx936": Prevent DTK from automatically adding multiple architectures resulting in --offload-arch leaked to g++HIP_USE_DEVICE_SORT=OFF: Avoid linking hip::device, prevent -xhip flag is passed to g++| Metric | GPU Package + HIP | Kokkos + HIP | Kokkos Advantages |
|---|---|---|---|
| Total time taken | 3.47 seconds | 0.55 seconds | 6.3× |
| timesteps/s | 288.1 | 1804.1 | 6.3× |
| Matom-step/s | 37.8 | 236.5 | 6.3× |
| Metric | GPU Package + HIP | Kokkos + HIP | Kokkos Advantages |
|---|---|---|---|
| Total time taken | 14.73 seconds | 3.20 seconds | 4.6× |
| timesteps/s | 67.9 | 312.5 | 4.6× |
| Matom-step/s | 71.2 | 327.7 | 4.6× |
fix gpu command to manage resourcesInspection of the DTK toolchain confirms that this build uses the Hygon native compiler:
# hipcc ultimately invokes Hygon's native dcc compiler
hipcc → <DTK_ROOT>/dcc/bin/dcc
↓
Hygon ISA bitcode: oclc_isa_version_936.bc
Kokkos' AMD_GFX936 is an architecture label. Compilation uses Hygon's dcc compiler and Hygon ISA bitcode, so the resulting code is compiled natively for the DCU.
Hotspots in the 1,048,576-atom dynamic LJ benchmark were analyzed with DTK 25.04.2 hipprof --hip-trace --stats. of the 1,048,576-atom dynamic LJ benchmark.
| Kernel | Number of calls | Total time | proportion |
|---|---|---|---|
| LJ pair force | 5,194 | 12.877 s | 64.92% |
| full neighbor-list build | 263 | 5.523 s | 27.85% |
| fused NVE integrate | 5,194 | 1.091 s | 5.50% |
| all remaining kernels | - | 0.344 s | 1.74% |
Pair and Neighbor account for 92.8% of device time and are the only worthwhile optimization targets.
PMC counters were collected for the Pair and Neighbor hotspots:
| Metric | LJ pair force | Neighbor build |
|---|---|---|
| VALU command | 70,299,737 | 433,198,843 |
| VMEM read command | 9,440,574 | 19,685,740 |
| VMEM write command | 65,536 | 9,601,271 |
| LDS command | 0 | 79,223,620 |
| LDS wait command | 0 | 14,343,770 |
| L2 hit rate | 95.10% | 83.77% |
| TCP data-stall cycle | 3,537,466 | 467,603,303 |
Pair does not use LDS and has a high L2 hit rate; Neighbor has a large number of LDS and TCP stall activities, which is the focus of the next optimization step.
tests workgroup 64/128/256/512/1024, staggered operation every three rounds:
| Workgroup | mean (step/s) | relative to 1024 |
|---|---|---|
| 64 | 228.706 | -8.58% |
| 128 | 230.307 | -7.94% |
| 256 | 233.920 | -6.50% |
| 512 | 241.857 | -3.32% |
| 1024 | 250.173 | Baseline |
improves monotonically as the workgroup increases, and the value of 1024 automatically selected by Kokkos is already the best value in the test. Experimental code has been rolled back.
Runtime options neigh/transpose on, no need to recompile, three rounds of interleaved verification:
| Mode | Run 1 | Run 2 | Run 3 | mean | Improve |
|---|---|---|---|---|---|
| transpose on | 284.280 | 284.007 | 284.166 | 284.151 | +13.61% |
| transpose off | 249.724 | 250.186 | 250.375 | 250.095 | Baseline |
changes the memory layout of the neighbor list on the GPU (row major → column major) and optimizes the cache access mode of the pair kernel. Attention transpose on is related to the system size: +13.6% for 1M atoms, but slightly slower -1.2% for 32K atoms.
tests the number of bins per team in the neighbor build (BPT=1/2/4/8), three rounds of staggered runs:
| BPT | mean (step/s) | relative to the default |
|---|---|---|
| 1 | 260.742 | +4.31% |
| 2 (default) | 249.974 | Baseline |
| 4 | 258.720 | +3.50% |
| 8 | 260.992 | +4.41% |
The default BPT=2 performs worst among the four configurations. Setting BPT=1 improves performance by 4.31% with a one-line change (const int factor = 1;); this setting is retained in the source.
The original report equates one DCU card with a 160-core CPU, based on results from different system sizes:
| Platform | Atomic number | timesteps/s | Matom-step/s |
|---|---|---|---|
| CPU (16 OMP threads) | 32,000 | 45.37 | 1.452 |
| DCU (original) | 256,000 | 912.28 | 233.544 |
The normalized throughput ratio is 233.544 / 1.452 = 160.8×, but the CPU and DCU results use different atom counts. The larger DCU case achieves higher GPU utilization, so this cross-scale ratio overstates the speedup for a like-for-like comparison.
| Platform | step/s | Matom-step/s | acceleration ratio |
|---|---|---|---|
| CPU (16 OMP threads) | 52.92 | 1.693 | 1× |
| DCU (1× C-3000 BW) | 3,206.09 | 102.595 | 60.6× |
At the same scale, one DCU card is about 60× faster, rather than 160×. Follow-up tests measured GPU utilization of 20–41%, depending on the pair style; CPU-side framework overhead was the main bottleneck.
| Optimization | Benchmark | Profit | Principle | Price |
|---|---|---|---|---|
| neigh/transpose on | LJ 1M | +13.61% | Neighbor list row-major → column-major, optimized pair kernel cache access mode | Runtime option, zero cost |
| neigh/transpose on | EAM 1M | +3.18% | Same as above, but EAM calculation is more intensive and memory access optimization has diminishing returns | Same as above |
| Neighbor BPT=1 | LJ 1M | +4.31% | Each team processes 1 bin (original default 2), halving LDS usage → higher occupancy | One line of source code modification |
| Neighbor BPT=1 | EAM 1M | +0.60% | Neighbor only accounts for 4.6% of the total EAM time, and there is very little room for optimization | Same as above |
| Optimization | Profit | Cause Analysis | Processing |
|---|---|---|---|
| Pair workgroup 64 | -8.58% | The smaller the workgroup, the fewer the number of waves per CU, and the degree of parallelism decreases | code has been rolled back |
| Pair workgroup 128 | -7.94% | performance increases monotonically with workgroup, Kokkos automatically selects 1024 Correct | code has been rolled back |
| Pair workgroup 256 | -6.50% | ||
| Pair workgroup 512 | -3.32% | ||
| DTK 25.04.4 upgrade | -0.23% | No performance difference between compiler/runtime version iterations | does not upgrade |
| DTK 26.04 upgrade | -0.11% | ||
| fix nve/kk mask fast path | -0.12% | Branch prediction fails, additional condition judgment increases overhead | code has been rolled back |
| fix nve/kk single type quality | -0.49% | Type checking overhead exceeds computational simplification benefits | code has been rolled back |
| nbin/atoms/per/bin scan | Unable to test | requires neigh/thread on, which will change the pair algorithm and confuse comparison | Give up |
| 2 GPU (no GPU-aware) | 0.42× | Comm accounts for 42%, data is copied back and forth between CPU and GPU | Upgrade to GPU-aware |
| 2 GPU (GPU-aware, LJ) | 0.71-0.79× | Comm down to 8-30%, but LJ calculation is too light to dilute communication | Use EAM |
| 2 GPU (GPU-aware, EAM 4M) | 0.78× | EAM is more computationally intensive, the 2/1 ratio has increased from 0.78 to 0.87, the trend is improving but has not exceeded 1.0 | Need larger scale |
| 2 GPU (GPU-aware, EAM 8M) | 0.87× |
LAMMPS multi-GPU requires MPI support. There is an ABI mismatch between the OpenMPI that comes with the system (compiled with GCC 8.5.0) and the current GCC 9.3.0, resulting in MPI_Finalize when double free crash. In addition, non-GPU-aware MPI requires CPU↔GPU data copy during each communication step, which results in huge communication overhead.
| Component | Version | Source |
|---|---|---|
| GCC | 9.3.0 | module load compiler/gcc/9.3.0 |
| DTK | 25.04.2 | module load compiler/dtk/25.04.2 |
| UCX | 1.18.0 | module load mpi/ucx/1.18.0/dtk-25.04/shca |
| OpenMPI | 5.0.7 | source code compilation |
# Download and build OpenMPI 5.0.7 wget https://download.open-mpi.org/release/open-mpi/v5.0/openmpi-5.0.7.tar.gz tar xzf openmpi-5.0.7.tar.gz && cd openmpi-5.0.7 # Load dependencies module purge module load compiler/gcc/9.3.0 module load compiler/dtk/25.04.2 module load <UCX_MODULE> # Configure UCX and ROCm GPU support ./configure --prefix=$HOME/softwares/mpi/install \ --enable-mpi-fortran=no \ --with-ucx=<UCX_ROOT> \ --with-rocm=<DTK_ROOT> make -j32 && make install
| Component | Description | Verification command |
|---|---|---|
accelerator: rocm | GPU direct memory access, avoiding CPU↔GPU copy | ompi_info --all | grep rocm |
pml: ucx | UCX point-to-point communication layer, supports InfiniBand/RoCE | ompi_info --all | grep 'pml: ucx' |
osc: ucx | UCX Unilateral communication (RDMA) | ompi_info --all | grep 'osc: ucx' |
mpiext: rocm | ROCm MPI extension | ompi_info --all | grep mpiext |
# Build LAMMPS with the locally built MPI
cmake -S lammps-22Jul2025/cmake -B build-mpi \
-C cmake/presets/hygon-dcu-gfx936.cmake \
-D CMAKE_CXX_COMPILER=$DTKROOT/bin/hipcc \
-D BUILD_MPI=ON \
-D MPI_CXX_COMPILER=$HOME/softwares/mpi/install/bin/mpicxx \
-D MPI_CXX_INCLUDE_PATH=$HOME/softwares/mpi/install/include \
-D MPI_CXX_LIBRARIES=$HOME/softwares/mpi/install/lib/libmpi.so
# Key Slurm job settings module load <UCX_MODULE> export PATH=$HOME/softwares/mpi/install/bin:$PATH export LD_LIBRARY_PATH=$HOME/softwares/mpi/install/lib:$LD_LIBRARY_PATH # Enable GPU-aware communication with -pk kokkos gpu/aware on mpirun -np 2 lmp -in input -k on g 1 -sf kk \ -pk kokkos neigh/transpose on gpu/aware on
| MPI version | 2 GPU step/s | Comm | shutdown | GPU-aware |
|---|---|---|---|---|
| System MPI (GCC 8.5, no UCX) | 104.65 | 42% | crash | No |
| Self-compiled MPI (GCC 9.3, no UCX) | 104.65 | 42% | ✅ | No |
| GPU-aware MPI (GCC 9.3 + UCX+ROCm) | 176.95 | 40% | ✅ | Yes |
GPU-aware MPI improves two-card performance by 69% (105 → 177 steps/s) and reduces communication time from 20.2 s to 11.4 s (−43%). The UCX + ROCm path avoids copying data through the CPU.
DTK 25.04.2 does not include UCX. The tested <UCX_MODULE> provides UCX components with ROCm transport support. Building OpenMPI 5.0.7 with --with-ucx and --with-rocm links these components and enables GPU-aware communication.
UCX module path:<UCX_ROOT>
ROCm path:<DTK_ROOT>
Platform B lacks the static library libhsakmt.a, so the UCX ROCm transport cannot be built. OpenMPI can still recognize device pointers through its ROCm accelerator extension and use shared memory for communication among cards on the same node:
./configure --prefix=<USER_HOME>/softwares/mpi/ompi-install \ --enable-mpi-fortran=no \ --with-rocm=<DTK_ROOT> \ --with-slurm export UCX_TLS=sm,self mpirun --mca pml ob1 --mca btl vader,self \ -np <GPU_COUNT> lmp -in input \ -k on g 1 -sf kk -pk kokkos gpu/aware on
This build exposes accelerator: rocm and mpiext: rocm. It was used for intra-node peer-to-peer tests with one to four cards; it does not validate cross-node RDMA.
Within the range of hardware, software versions and workloads described in this report,KOKKOS + HIP + neigh/transpose on achieved the best overall performance.
| Pair Style | Mode | 1 GPU | 2 GPU | 4 GPU | 2/1 | 4/1 |
|---|---|---|---|---|---|---|
| EAM 1M | KOKKOS | 76.5 | 57.1 | 45.2 | 0.75 | 0.59 |
| EAM 4M | KOKKOS | 21.2 | 16.5 | — | 0.78 | — |
| EAM 8M | KOKKOS | 10.2 | 8.9 | — | 0.87 | — |
| EAM 16M | KOKKOS | 5.2 | 4.7 | — | 0.90 | — |
| LJ 1M | KOKKOS | 277.7 | 198.3 | — | 0.71 | — |
| Tersoff 1.1M | KOKKOS | 195.0 | 123.2 | — | 0.63 | — |
| EAM 1M | GPU package | 39.6 | — | — | — | — |
EAM 2/1 ratio 0.78→0.87→0.90, trend improving but not breaking 1.0.GPU package has good scalability (2/1=1.32) but single card is slow (~1.8× KOKKOS).
| vs. | Original report | After correction | Question |
|---|---|---|---|
| acceleration ratio | 160× | 60× | Original cross-scale comparison (32K CPU vs 256K DCU) unfair |
| Acceleration trend | — | As the number of atoms increases | 32K: 60× → 1M: ~172× (estimate) |
| Atomic number | step/s | Matom-step/s | Pair% | memory |
|---|---|---|---|---|
| 1M | 72.5 | 76.0 | 36.7% | 0.16GB |
| 4M | 21.2 | 85.0 | 32.8% | 0.64GB |
| 8M | 10.2 | 81.9 | 33.8% | 1.2GB |
| 16M | 5.2 | 82.6 | 32.2% | 2.4GB |
| 25M | 3.2 | 80.2 | 32.0% | 3.8GB |
| 32M | — | — | — | OOM |
Throughput is stable at 80–85 Matom-step/s with near-linear scaling. The 32M-atom case runs out of memory because the neighbor list initially allocates 2,000 entries per atom, exceeding 64 GB. Memory capacity becomes the limit before compute throughput.
dihedral charmm/kk implements only HALF neighbor lists, while GPU mode defaults to neighflag=FULL. atom_style full together with lj/charmm/coul/long also forces FULL lists, creating a conflict. Running the post-rhodo case on the GPU with -pk kokkos neigh half matches the CPU result. The missing FULL-list implementation remains an upstream limitation.
| Category | Restrictions | Process/Affect |
|---|---|---|
| DTK 26.04 | host mode injection cuda_wrappers | Change lmp target to -xhip --offload-arch=gfx906 |
| DTK 26.04 | Missing libhsakmt.a | Unable to build UCX ROCm; use OpenMPI --with-rocm |
| HSA runtime | symbolic chain depends on /opt/hyhal | Explicitly specify the real library in the shared software area |
| Compilation | generates gfx906/926/928 three-architecture by default | Fixed AMDGPU_TARGETS=gfx906 |
| Hardware | gfx906 No FP8/BF16 support or MFMA matrix cores | This does not affect the molecular-dynamics tests in this report; results should not be extrapolated to matrix workloads |
| GPU memory | 16 GB per card; LJ runs are limited to about 16 million atoms | Larger systems require multiple cards or lower memory use |
| Scheduling | The test partition rejects requests for more than 8 CPU cores | The multi-card test used 8 CPU cores in total; the partition name is redacted |
--with-rocm and shared memory; Platform A uses the UCX+ROCm path.pml ob1, btl vader,self, and UCX_TLS=sm,self to avoid GID assertions in unsupported InfiniBand/UCX paths.| Contribution | Content | file | Status |
|---|---|---|---|
| gfx926/928/936/938 architecture support | Adds Hygon DCU architecture to Kokkos cmake and HIP core | kokkos_arch.cmake, KokkosCore_config.h.in, Kokkos_HIP_Instance.hpp, Kokkos_HIP_IsXnack.hpp | prototype has been verified, waiting for upstream review |
| wavefront 64 support | gfx936 uses 64-wide wavefront, same as gfx906 | Kokkos_HIP_Instance.hpp | Local implementation verified |
| System allocated access attribute | Hygon DCU does not support direct access to system memory | Kokkos_HIP_IsXnack.hpp | Local implementation verified |
| Contribute | Content | Status |
|---|---|---|
| build preset | cmake/presets/hygon-dcu-gfx936.cmake: Reproducible Hygon DCU build configuration | To be submitted |
| Performance Benchmark | KOKKOS vs GPU package comparison data, multi-GPU scaling analysis | available for document reference |
| Recommended configuration | neigh/transpose on has +3~14% improvement over gfx936 | Runtime options, no code required |
| Contribution | Content | Status |
|---|---|---|
| gfx936 performance data | Single/multi-card performance of LJ, EAM, Tersoff, ReaxFF on 80 CU | available for document reference |
| GPU-aware MPI solution | OpenMPI 5.0.7 + UCX 1.18.0 DTK + ROCm compilation method | Documented |
| Verification method | PMC analysis tools, HIP trace workflow, dynamic benchmark | tools/hygon/ |