opensource_community
Complete technical report · Baseline record 2026-08-21 · Latest results updated to 2026-08-26 · Redacted public edition

Porting and Debugging vLLM 0.27.x for DeepSeek V4 Flash FP4 on Hygon DCU gfx936

2026-08-26 Core results:vLLM 0.27.x has been successfully ported to Hygon DCU gfx936. Following targeted DTK adaptations, the V1 Engine completed end-to-end inference with the actual DeepSeek V4 Flash FP4 weights at TP=4 and TP=8. The mixed-precision path—including Sparse MLA, MXFP4 MoE, and continuous decoding—runs end to end. Work has moved from basic functional porting to stability and performance engineering.

0. Supplementary verification results (2026-08-24 to 08-26)

0.1 From functional validation to a stable performance baseline

0.2 Performance and profile

ConfigurationresultsConclusion
TP=4, single request, fixed 32 tokensAbout 0.99 tok/sCurrent recommended baseline
TP=8, single request, fixed 32 tokensAbout 0.68–0.81 tok/sCommunication overhead offsets additional computing resources
TP=4, batch=8 historical measurementAbout 2.5 aggregate tok/sFailed correctness checks; excluded from official performance results

The clean profile attributes about 30% of steady-state GPU time to the shared MoE forward pass and about 14% to tensor-parallel communication. Many smaller operations (to, copy, sum, add, zeros, and repeat_interleave) also contribute. Disabling routing Q1 increased fixed-generation time from about 67.9 s to 101.3 s, so Q1 must remain enabled. The shared-expert auxiliary-stream change regressed end-to-end performance and was abandoned.

0.3 Batch correctness limits

Batch sizes 1–7 do not show the structural BOS failure. At batch size 8, long decoding can produce a nondeterministic BOS token, degraded output, or NaNs. The issue also occurs with the PyTorch Sparse reader; MXFP4 tests for M=1–9 pass bitwise, and NaNs have been observed before the sampler. The leading areas to investigate are asynchronous streams, workspace reuse, and kernel races—not simply the Triton reader or M=8 GEMM.

The new Q1 cast-elision and grouped-batch MoE paths remain disabled by default. Before either is recommended, it must pass checks against real operator outputs, end-to-end correctness, and a fixed-workload throughput A/B comparison. The latest throughput scan could not run because of a scheduling-account error, so it produced no new performance evidence.

0.4 Stable configuration

DSV4_USE_TRITON_SPARSE=1
DSV4_MOE_ROUTING_Q1=1
DSV4_FP8_WEIGHT_CACHE=0
DSV4_FP8_TRITON_GEMM=0
DSV4_ROCM_SHARED_EXPERTS_STREAM=0
DSV4_GREEDY_LOCAL_ARGMAX=0
DSV4_MOE_Q1_PREPARE=0
VLLM_ROCM_USE_SKINNY_GEMM=0
DSV4_MOE_BACKEND=emulation

Baseline report date: 2026-08-21; latest update: 2026-08-26
Final status: End-to-end inference passed, with coherent output during continuous decoding
Baseline verification job: J17 — COMPLETED / ExitCode 0:0
Target platform: eight Hygon DCU gfx936 cards; DTK 26.04
Inference framework: vLLM 0.27.2.dev0+g6e448d0ea.d20260819
Model: DeepSeek-V4-Flash-0731

1. Executive Summary

The goal was not merely to load the model weights, but to run DeepSeek V4 Flash FP4 end to end through vLLM on Hygon DCU gfx936:

  1. Initialize tensor parallelism across eight cards;
  2. Load approximately 155 GiB of model weights across 48 shards;
  3. Run FP4 MoE expert weights through the MXFP4 emulation path;
  4. Dequantize FP8 backbone weights in software on gfx936;
  5. Use the fp8_ds_mla KV cache for DeepSeek V4 Sparse MLA;
  6. Continue decoding after prefill;
  7. Produce semantically coherent text;
  8. Complete the Slurm workload with an actual exit code of 0.

Output from the successful run:

TOKEN_IDS: [223, 30594, 9790, 303, 70037, 12203, 10706, 804, 1175,
            10539, 637, 53897, 96240, 2996, 84998, 666, 3118, 637,
            58368, 621, 666, phase, 53091, 4374, 1465, 223, 1959,
            17585, 42108, 303, 56463, 10162]

TOKEN_DECODE:
[' ', '你好', '呀', ' ,', '很高兴', '在这里', '遇到', '你', '!',
 '让我', '来', '做个', '自我介绍', '吧', '~\n\n', '**', '初', '来',
 '乍', '到', '**', ':', 'Deep', 'Se', 'ek', ' ', '这个', '版本',
 '的我', ' ,', '确实是', '有点']

OUTPUT_REPR:
' 你好呀 ,很高兴在这里遇到你!让我来做个自我介绍吧~\n\n
 **初来乍到**:DeepSeek 这个版本的我 ,确实是有点'

DSV4_VERDICT: COHERENT

Final job status:

JobID|State|ExitCode|Elapsed|MaxRSS
J17|COMPLETED|0:0|00:05:15|
J17.batch|COMPLETED|0:0|00:05:15|65899600K
J17.extern|COMPLETED|0:0|00:05:15|0

The root cause was not the FP4 MoE numerical format, chat template, skinny GEMM, or tokenizer. It was the gfx936 Sparse MLA PyTorch decode fallback misreading the fp8_ds_mla KV cache.

KV cache is:

[num_blocks, block_size, 584]

But each block is not 64 independent 584-byte tokens arranged in sequence in physical memory, but:

[64 × 576-byte token data][64 × 8-byte UE8M0 scales]

The original reader incorrectly assumed that each token occupied 584 contiguous bytes. As a result, it:

After the fix, the reader and C++ writer use matching address calculations:

data  = block_base + token_pos * 576
scale = block_base + block_size * 576 + token_pos * 8

Verification shows:

scale_bytes = [118, 119, 119, 119, 118, 119, 118, 0]
NoPE finite = True
RoPE finite = True
attention logits finite = True
attention probs finite = True

Together, these checks form a complete evidence chain from physical addresses and quantization scales through KV-cache values, attention logits, and token IDs to the final natural-language output.


2. Goals, scope, and success criteria

2.1 Core Objectives

The core goal has always been:

runs through DeepSeek V4 Flash FP4 inference through vLLM on Hygon DCU gfx936.

The "FP4" here specifically refers to the model's MoE expert using MXFP4 weights. Not all tensors of the model are FP4; the actual checkpoint also contains:

Therefore, the real goal is to run through the mixed precision execution chain rather than verifying an FP4 GEMM alone.

2.2 Outcomes that do not constitute completion

The following states do not count as a successful end-to-end run:

2.3 Final success criteria

The end-to-end success criteria used in this round are:


3. Hardware and software environment

3.1 Computing resources

Measured environment:

Project value
accelerator Hygon DCU
Architecture gfx936:sramecc+:xnack-
Number of cards 8
GPU memory per card (as reported by PyTorch) Approximately 68.7 GB decimal, which is approximately 63.98 GiB
Slurm Partition 内部 DCU 分区
CPU --cpus-per-task=64
working memory --mem=240gb
Maximum test time 02:00:00

3.2 Software Stack

Component version or path
DTK 26.04
GCC module 12.2.0
Python 3.12 ,vLLM venv
vLLM 0.27.2.dev0+g6e448d0ea.d20260819
model DeepSeek-V4-Flash-0731
vLLM source code <VLLM_SOURCE>
venv <VLLM_ENV>
Test Catalog <TEST_WORKSPACE>

3.3 Model and quantization configuration

Read from the running log:

quant_config:
{
  'activation_scheme': 'dynamic',
  'fmt': 'e4m3',
  'quant_method': 'fp8',
  'scale_fmt': 'ue8m0',
  'weight_block_size': [128, 128]
}

vLLM runtime confirmation:

DeepSeek V4 expert_dtype resolved to 'fp4'
Using 'EMULATION' Mxfp4 MoE backend.
Using DeepSeek's fp8_ds_mla KV cache format.
Using FP8 indexer cache for Lightning Indexer.

3.4 Runtime environment variables

export VLLM_WORKER_MULTIPROC_METHOD=spawn
export TORCHINDUCTOR_USE_STATIC_CUDA_LAUNCHER=0
export DSV4_SHIM_DEBUG=1
export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1

export DSV4_TP=8
export DSV4_MAXLEN=512
export DSV4_MEMUTIL=0.90
export DSV4_KV_CACHE_DTYPE=fp8
export DSV4_MOE_BACKEND=emulation

export VLLM_ROCM_USE_SKINNY_GEMM=0

Also used during debugging:

export DSV4_SPARSE_DEBUG=1

4. Implementation approach

The end-to-end execution chain can be simplified to:

Model config / tokenizer
        │
        ▼
FP8 backbone weights ──software dequantization──► BF16 GEMM
        │
        ├──► Q / KV / indexer / compressor
        │
        ▼
FP4 MoE expert ──MXFP4 emulation──► MoE output
        │
        ▼
DeepSeek V4 Sparse MLA
        │
        ├── prefill
        ├── SWA cache write(C++ fused writer)
        ├── compressed/top-k cache
        └── decode fallback(gfx936 PyTorch reference)
                │
                ▼
        logits → softmax → attention output
                │
                ▼
        LM head → next token

The investigation uncovered four independent compatibility and correctness issues:

  1. Triton / DAS Triton API and lowering compatibility issues;
  2. A device-side assertion in gfx936 skinny GEMM;
  3. Input format misjudgment caused by prompt/chat template;
  4. The physical layout of the Sparse MLA KV cache reader is wrong.

Category 4 is the final root cause of "the first token is normal, and subsequent tokens are all BOS".


5. Diagnostic methods and test probes

The diagnostic probe reports progress by stage:

### STAGE: env
### STAGE: config
### STAGE: vllm import + registry
### STAGE: engine init (the real test)
### STAGE: generate

Answer each stage separately:

Stage Questions to answer
env Can PyTorch detect all eight gfx936 devices, and are their memory sizes reported correctly?
config Can the model architecture, layer count, expert configuration, and quantization settings be read correctly?
registry Is DeepSeek V4 registered with vLLM?
engine init Did weight loading, tensor parallelism, MoE setup, KV-cache allocation, and warm-up complete?
generate Do prefill and continuous decoding produce valid tokens?

Print during generation:

TOKEN_IDS
FINISH_REASON
TOKEN_DECODE
OUTPUT_REPR
OUTPUT
DSV4_VERDICT

Slurm script explicitly saves Python return code:

python load_dsv4.py > "$RAW" 2>&1
rc=$?
echo "PY_EXIT=$rc"
exit "$rc"

So the Slurm status, Python verdict and raw logs in the report can be cross-validated.


6. Validation timeline and job records

6.1 Summary table

Job Status Time reaches level Main results
J01 FAILED 1 06:18 engine init Triton lowering PassManager::run failed
J02 FAILED 1 06:18 engine init Triton softmax(dim=...) API is not compatible with
J03 FAILED 1 06:41 engine init Similar softmax lowering problem
J04 FAILED 1 06:16 engine init Similar softmax lowering problem
J05 FAILED 1 06:45 engine init fallback still triggers softmax API issue
J06 FAILED 1 06:43 generate ENGINE_READY;skinny GEMM device assertion
J07 FAILED 2 06:49 generate Disable skinny GEMM; no hard crashes, output degradation
J08 COMPLETED 0 07:11 generate [32,0,0,...]; Probe false alarm COHERENT
J09 FAILED 2 03:32 generate Still [0,0,...]
J10 CANCELLED 00:16 Start The old script submitted by mistake, proactively canceled
J11 FAILED 2 07:13 generate chat template repair; first token 你好, follow-up BOS
J12 FAILED 1 01:15 engine init auto KV dtype was replaced by fp8_ds_mla layout reject
J13 FAILED 2 03:05 generate confirms again that the first token is random and subsequent BOS
J14 FAILED 2 03:06 generate diagnostics are consumed by warmup and the real decode
J15 FAILED 2 04:51 generate captures the real decode: cache K/attention full NaN
J16 FAILED 2 07:58 generate The first stride repair is insufficient and still NaN
J17 COMPLETED 0 05:15 end-to-end The physical layout was corrected, and the output is consistent

Interpretation:

6.2 J01: Triton lowering blocks engine startup

RuntimeError: PassManager::run failed
RuntimeError: Engine core initialization failed.

At this stage:

6.3 J02–J05: Resolve softmax API incompatibility

Stable reproduction of multi-round operations:

triton.compiler.errors.CompilationError
TypeError("softmax() got an unexpected keyword argument 'dim'")

This is the difference between the upstream code and the current DAS Triton API. At this stage, by adjusting fallback and avoiding incompatible Triton paths, the failure position is gradually advanced to the real engine warmup.

Engineering lesson: Avoid repeatedly patching isolated Triton syntax such as tt.dot_scaled. For operations that cannot be lowered on gfx936, provide a testable PyTorch/BF16 reference fallback.

6.4 J06: Engine startup succeeds; first device-side error

Key developments:

Model loading took 19.99 GiB memory and 205.150379 seconds
GPU KV cache size: 85,259 tokens
ENGINE_READY in 335.4s

The first device-side error during generation:

csrc/rocm/skinny_gemms.hip:580
wvSplitK_hf_sml_
Device-side assertion `false' failed.

This is a true device-side error followed by a VMFault.

Workaround:

export VLLM_ROCM_USE_SKINNY_GEMM=0

With this setting:

6.5 J07–J08: Crashes resolved, but output remains incorrect

J08:

TOKEN_IDS: [32, 0, 0, ..., 0]
TOKEN_DECODE:
['>', '<|begin▁of▁sentence|>', ...]
OUTPUT_REPR: '>'

The probe initially reported:

DSV4_VERDICT: COHERENT

The probe's success check only tested whether:

Because OUTPUT_REPR contained only one >, it passed the weak “nonempty and not repetitive” check, producing a false positive.

The acceptance check was strengthened to inspect all of the following:

6.6 J09: Add a SwiGLU fallback

The original gfx936 emulation fallback was:

act = F.silu(gate * 1.702) * up

This did not match the model configuration or the vLLM reference. It was changed to:

gate = torch.clamp(gate, max=gemm1_limit)
up = torch.clamp(up, min=-gemm1_limit, max=gemm1_limit)
act = F.silu(gate * gemm1_alpha) * up

and uses:

gemm1_alpha = 1.0
gemm1_limit = 10.0

Job result:

TOKEN_IDS: [0, 0, ..., 0]

Conclusion:

6.7 J11: Fix the chat template and isolate the decode issue

A plain prompt:

llm.generate(["你好 ,请介绍一下你自己 ."])

does not necessarily follow DeepSeek V4's custom conversation template. After switching to the chat API:

TOKEN_IDS: [30594, 0, 0, ..., 0]
TOKEN_DECODE:
['你好', '<|begin▁of▁sentence|>', ...]
OUTPUT_REPR: '你好'

This is very critical boundary evidence:

6.8 J12: Test BF16 and automatic KV-cache settings

The automatic KV-cache setting was tested:

DSV4_KV_CACHE_DTYPE=auto

Engine initialization rejected it explicitly:

AssertionError:
DeepseekV4 fp8_ds_mla layout only supports fp8 kv-cache, got auto

The current DeepSeek V4 ROCm Sparse MLA implementation is bound to the fp8_ds_mla layout. Setting the cache type to auto cannot bypass this constraint; the layout's read and write paths must be implemented correctly.

6.9 J13: Confirm repeatability

ENGINE_READY in 93.2s
TOKEN_IDS: [57297, 0, 0, ..., 0]
TOKEN_DECODE:
[' Modified', BOS, BOS, ...]

The first generated token varied between runs:

But the common pattern is always:

[prefill token, BOS, BOS, BOS, ...]

Therefore the problem cannot be interpreted as a fixed tokenizer mapping; the failure occurs in the decode calculation chain.

6.10 J14: Initial diagnostic probe misses the failing decode

First diagnostic output:

q shape = (8, 8, 512)
main_indptr = [0, 0, ..., 0]

This is the dummy invocation of engine warmup, not a real request. The one-time diagnostic flag is consumed by warmup, causing the real decode to not be printed.

Corrected trigger condition:

diag_once = (
    diag
    and num_queries <= 2
    and not getattr(fn, "DSV4_DIAG_ONCE", False)
)

num_queries=1, warmup's num_queries=8.

6.11 J15: Numerical evidence identifying the root cause

Metadata from the failing decode:

q shape          = (1, 8, 512)
main_cache shape = (35644, 64, 584)
main_cache stride= (1002240, 584, 1)
main_indices     = [128, 129, ..., 138, ...]
main_indptr      = [0, 11]
scale            = 0.04419417382415922
qfinite          = True

This proves:

After decoding, however, the cache contained:

scale_bytes:
[87, 245, 226, 232, 107, 107, 156, 110]

nope_finite = False
nope_max    = nan

rope_finite = True
rope_max    = 1.0135363467859984e+37

Attention diagnostics:

logits_finite = False
logits_min    = nan
logits_max    = nan
probs_finite  = False
probs_max     = nan

The complete failure chain is:

Q finite
  → cache decode error
  → NoPE NaN / RoPE 1e37
  → attention logits NaN
  → softmax NaN
  → hidden-state corruption
  → LM-head degradation
  → token 0 / BOS

6.12 J16: Why the first stride fix was insufficient

The initial fix made the cache contiguous:

cache.view(torch.uint8).contiguous()

That change was removed. The next attempt kept the original cache view and read it with:

raw[block, pos]

The result remained incorrect:

scale_bytes = [88, 245, 226, 232, ...]
nope_finite = False
rope_max    ≈ 1e37

This showed that the issue was not simply a lost stride(0) caused by .contiguous().

Comparison with the C++ writer showed that stride(1)=584 cannot be used by itself to locate individual tokens. The physical cache is block-packed rather than token-interleaved.

6.13 J17: Correct the physical cache layout

Final diagnosis:

scale_bytes:
[118, 119, 119, 119, 118, 119, 118, 0]

nope_finite = True
nope_max    = 1.875

rope_finite = True
rope_max    = 5.75 ~ 5.78125

Attention values were valid across tensor-parallel ranks. For example:

TP0:
logits_finite = True
logits_min    = -2.2683308124542236
logits_max    =  6.992861270904541
probs_finite  = True
probs_max     =  0.7818039655685425

TP2:
logits_finite = True
logits_min    = -3.1048312187194824
logits_max    =  4.642844200134277
probs_finite  = True
probs_max     =  0.5944044589996338

TP7:
logits_finite = True
logits_min    = -2.159546136856079
logits_max    =  6.589494228363037
probs_finite  = True
probs_max     =  0.9574916362phase

The final token no longer degrades and forms coherent Chinese.


7. Root cause: logical shape and physical KV-cache layout diverge

7.1 Logical tensor representation

The backend declares that each token occupies 584 bytes:

448 byte FP8 NoPE
128 byte BF16 RoPE
8 byte UE8M0 scales/padding

logical shape:

[num_blocks, block_size, 584]

This is easy for the reader to write:

token = cache[block, pos]
data = token[:576]
scale = token[576:584]

But this explanation is wrong.

7.2 Addresses written by the C++ writer

writer uses:

block_base =
    k_cache + block_idx * kv_block_stride;

token_fp8_ptr =
    block_base + pos_in_block * kTokenDataBytes;

token_bf16_ptr =
    token_fp8_ptr + kNopeDim;

token_scale_ptr =
    block_base
    + cache_block_size * kTokenDataBytes
    + pos_in_block * kScaleBytesPerToken;

Constant:

kNopeDim            = 448
kRopeDim            = 64 BF16 = 128 byte
kTokenDataBytes     = 576
kScaleBytesPerToken = 8
block_size          = 64

Therefore the single block physical layout:

offset 0
│
├── token 0 data:   576 B
├── token 1 data:   576 B
├── ...
├── token 63 data:  576 B
│
├── token 0 scale:  8 B
├── token 1 scale:  8 B
├── ...
└── token 63 scale: 8 B

Total valid data:

64 × 576 + 64 × 8
= 36,864 + 512
= 37,376 byte

while running observed:

main_cache.stride() = (1002240, 584, 1)

stride(0) is determined by the allocator/shared cache layout, which is much larger than the valid data of a single block. The reader must retain the real block stride; it cannot first .contiguous().

7.3 Two errors in the original reader

Old code:

cache_u8 = cache.view(torch.uint8).contiguous()
raw = cache_u8.reshape(cache.shape[0], block_size, -1)
token = raw[block, pos]
data = token[:, :576]
scale_bytes = token[:, 576:584]

Error 1:.contiguous() destroys physics stride(0).

Error 2: Even if stride is retained,token[:, 576:584] still assumes that scale follows the data of each token; the actual scale is concentrated at the end of the block.

7.4 Final reader

The final implementation first constructs a byte view that retains block stride:

raw = cache_u8.as_strided(
    (cache.shape[0], block_size * 584),
    (cache_u8.stride(0), 1),
)

Read data:

data_offsets = (
    pos[:, None] * 576
    + torch.arange(576, device=device)[None, :]
)
data = raw[block[:, None], data_offsets]

Reading scale:

scale_offsets = (
    block_size * 576
    + pos[:, None] * 8
    + torch.arange(8, device=device)[None, :]
)
scale_bytes = raw[block[:, None], scale_offsets]

UE8M0 decoding:

k_scales = torch.exp2(scale_bytes.float() - 127.0)

FP8 NoPE:

k_nope = data[:, :448].contiguous().view(fp8_dtype).float()
k_nope *= k_scales[:, :7].repeat_interleave(64, dim=1)

BF16 RoPE:

k_rope = (
    data[:, 448:576]
    .contiguous()
    .view(torch.bfloat16)
    .float()
)

7.5 Basis for the final scale

Final scale bytes:

[118, 119, 119, 119, 118, 119, 118, 0]

correspond to 7 64-element groups of 448 NoPE elements.

For example:

encoded 118 → scale = 2^(118 - 127) = 2^-9
encoded 119 → scale = 2^(119 - 127) = 2^-8

is padding, the value is 0, and does not participate in the dequantization of the 7 actual groups.

This is exactly the same as the writer’s 7 UE8M0 scale + 1 byte pad design.


8. Hypotheses tested and their limits

8.1 Hypothesis: the FP4 MoE implementation is faulty

The evidence does not support this.

But the SwiGLU parameter fix should still be retained as it fixes model formula consistency.

8.2 Hypothesis: the tokenizer or chat template is the sole cause

is only partially true.

8.3 Hypothesis: switching the KV cache to auto/BF16 avoids the issue

is not established.

Current DeepSeek V4 ROCm Sparse MLA layout mandatory requirement:

fp8_ds_mla

auto was explicitly rejected in engine init.

8.4 Hypothesis: removing contiguous() is sufficient

is not established.

.contiguous() does lose block stride, but the logical token stride also cannot represent the physical layout. Must also be fixed:

  1. block stride ;
  2. data/scale partition address.

8.5 Hypothesis: a correct first token proves overall numerical correctness

is not established.

The first generated token is produced during prefill; single-token decoding begins only when the next token is requested. The fault occurs at the prefill-to-decode transition.

Therefore end-to-end testing requires at least:

max_tokens >= 2

A safer bet would be 16 or 32.

8.6 Hypothesis: a COMPLETED Slurm job proves model correctness

is not established.

J08 is a typical counterexample:

State = COMPLETED
ExitCode = 0
TOKEN_IDS = [32, 0, 0, ...]

framework does not throw an exception, which does not mean that the value is correct. Compatibility report must also be checked:


9. Implemented changes

9.1 gfx936 compatibility shim

sitecustomize.py is automatically applied after worker spawn to make up for the current Triton/vLLM API differences and avoid premature import of Triton at startup.

include:

These are the infrastructure that the engine can use to start and complete weight loading.

9.2 Software dequantization for the FP8 backbone

gfx936 does not support native FP8 as expected by vLLM _scaled_mm path, so FP8 linear uses BF16 software inverse quantization fallback.

This path sacrifices performance but establishes a baseline for correctness.

9.3 MXFP4 MoE emulation

DSV4_MOE_BACKEND=emulation

weights remain in FP4 checkpoint format and are dequantized and calculated on demand by the emulation expert path. This path is used to first prove that the model functions correctly.

9.4 SwiGLU parameter alignment

Fixed the error 1.702 formula with the one consistent with the model configuration:

alpha = 1.0
limit = 10.0

9.5 Disable gfx936 skinny GEMM

VLLM_ROCM_USE_SKINNY_GEMM=0

Avoid wvSplitK_hf_sml_ .

9.6 Sparse MLA decode fallback

gfx936 cannot stabilize lower, so PyTorch reference decode is used.

The final critical fix is located at:

vllm/v1/attention/ops/rocm_aiter_mla_sparse.py

The repair content is fp8_ds_mla block-packed cache.


10. Verification results

10.1 Engine initialization

ENGINE_READY in 228.9s

Across different nodes, cache conditions, and JIT states, initialization previously ranged from about 93 to 338 seconds. A single engine-ready measurement therefore cannot establish steady-state performance.

10.2 KV-cache values

All eight tensor-parallel workers reported the same scale bytes:

[118, 119, 119, 119, 118, 119, 118, 0]

and:

NoPE finite = True
NoPE abs max = 1.875
RoPE finite = True
RoPE abs max ≈ 5.75–5.78125

10.3 Attention values

Across all tensor-parallel ranks:

logits_finite = True
probs_finite  = True

Logits and maximum softmax probabilities differ slightly across ranks, as expected because each rank holds different attention-head shards.

10.4 Output quality

The 32-token response contains no repeated BOS tokens. Its Chinese text is coherent and includes natural punctuation and Markdown formatting:

你好呀 ,很高兴在这里遇到你!让我来做个自我介绍吧~

**初来乍到**:DeepSeek 这个版本的我 ,确实是有点

10.5 Resources

MaxRSS = 65899600K

This is the batch step's CPU resident memory. The GPU-side log reports:

Model loading took approximately 19.99 GiB per worker/device
GPU KV cache size approximately 85K tokens

These values depend on the node, JIT state, allocator, and cache configuration. They document this functional-validation run and should not be treated as general performance figures.


11. Log analysis: distinguish noise from actionable errors

11.1 Usage-report JSON error

The following message appeared in several logs:

Exception in thread Thread-1 (_report_usage_worker)
json.decoder.JSONDecodeError

The warning originates in the usage-report background thread, not in engine startup or model computation. Similar noise also appears in runs that complete successfully.

11.2 libcudart assertion during ROCm shutdown

Failed jobs commonly emit this message during cleanup:

AssertionError:
libcudart is not loaded in the current process,
try setting VLLM_CUDART_SO_PATH

This is a ROCm compatibility issue with the vLLM shutdown path trying to import the CUDA allocator. It appears after generating results and probe verdicts, and should not be misjudged as the root cause of this inference failure.

11.3 How to read FIRST BLOCKER

Identify the first blocker in this order:

  1. Find the earliest workload traceback in the raw log;
  2. Confirm whether it occurs in the main execution path;
  3. distinguish background-thread and shutdown messages from errors in the main inference path;
  4. compare ENGINE_READY, TOKEN_IDS, and the probe verdict;
  5. Finally check the Slurm state and exit code.

12. Reproduction experiment

12.1 Slurm resources

#SBATCH -p <DCU_PARTITION>
#SBATCH --gres=dcu:8
#SBATCH --cpus-per-task=64
#SBATCH --mem=240gb
#SBATCH --time=02:00:00

12.2 Environment

module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VLLM_ENV>/bin/activate

export LD_LIBRARY_PATH=<DTK_ROOT>/gcvm/lib:\
<DTK_ROOT>/lib:\
<DTK_ROOT>/lib64:\
<DTK_ROOT>/comgr/lib64:\
$LD_LIBRARY_PATH

export PYTHONPATH=<TEST_WORKSPACE>:\
<TRITON_KERNELS>:\
$PYTHONPATH

12.3 vLLM parameters

llm = LLM(
    model=MODEL,
    tensor_parallel_size=8,
    max_model_len=512,
    kv_cache_dtype="fp8",
    moe_backend="emulation",
    enforce_eager=True,
    gpu_memory_utilization=0.90,
    trust_remote_code=True,
)

Dialog input should use the model chat template:

llm.chat(
    [[
        {
            "role": "user",
            "content": "你好 ,请介绍一下你自己 .",
        }
    ]],
    SamplingParams(
        max_tokens=32,
        temperature=0,
        ignore_eos=True,
    ),
)

12.4 Result checks

Must check:

ENGINE_READY
TOKEN_IDS
TOKEN_DECODE
OUTPUT_REPR
DSV4_VERDICT
PY_EXIT
sacct State / ExitCode

should not just grep OUTPUT:because special tokens such as BOS may be omitted from the final text by the tokenizer.


13. Code version and supporting files

Local warehouse:

file Purpose
probes/load_dsv4.py Phased end-to-end loading and generation of probes
probes/load_dsv4.slurm Basic Slurm Script
probes/sitecustomize.py spawn worker automatically compatible with shim
probes/patch_sparse_diag_min.py Sparse MLA Q/cache/attention numerical diagnosis
probes/patch_sparse_physical_layout.py Final physical cache layout repair script
dsv4-vllm-handoff.md Earlier torch/vLLM/mmac porting history
dsv4-dcu-record.html Project process record

Remote key source code:

<VLLM_SOURCE>/
  vllm/v1/attention/ops/rocm_aiter_mla_sparse.py

The remote end retains pre-repair backup and diagnostic scripts for easy comparison.


14. Findings and scope

14.1 Proven

14.2 Not yet proven

Therefore the current expression should be:

FP4 functional link and continuous decode have been run through and have a correctness baseline; production-level stability, accuracy and performance acceptance still require subsequent testing.

15. Recommended follow-up work

15.1 Priority 1: Make regression tests permanent

Adds at least the following regressions:

  1. single request, 32 tokens, Chinese chat;
  2. single request, 128 tokens;
  3. Two rounds of dialogue;
  4. 8 different prompts;
  5. Two concurrent requests;
  6. prompt length crosses block boundary;
  7. SWA window border;
  8. C4/C128 compressor boundary;
  9. prefix caching switch A/B;
  10. runs 50 times to check for occasional NaNs.

Must check every time:

all token IDs
all special tokens
finite checks
Slurm exit code
VMFault/device assertion

15.2 Priority 2: Turn diagnostics into test assertions

Upgrade one-time printing to optional assertion:

assert torch.isfinite(k_nope).all()
assert torch.isfinite(k_rope).all()
assert torch.isfinite(logits).all()
assert torch.isfinite(probs).all()

Only in DSV4_SPARSE_VALIDATE=1 is used to avoid loss of production performance.

15.3 Priority 3: Compare numerical accuracy

Recommended layering:

  1. cache writer/reader round-trip ;
  2. single token attention output;
  3. single layer hidden state cosine;
  4. Hidden state of each layer;
  5. next-token logits top-k ;
  6. greedy token sequence ;
  7. small data set perplexity.

The reference baseline can be:

15.4 Priority 4: Performance optimization

The current PyTorch decode fallback will produce:

It is suitable as a correctness reference, not suitable as a final high-performance implementation.

The next step should be to port the correct physical address formula to:

must retain this round of reference implementation as an oracle before optimization.

15.5 Priority 5: Remove temporary debugging code

Before production:


16. Engineering lessons

16.1 Tensor shape does not determine physical layout

Quantized caches often use logical shapes that "conveniently express capacity", but writers may use completely different physical layouts for vectorization, TMA, coalescing, or scale partitions.

The reader must use the writer's address formula as the source of truth.

16.2 Instrument prefill and decode separately

warmup, prefill and decode may be used differently:

The "first function call" is usually just a warmup and does not represent a real request.

16.3 A correct first token is a useful diagnostic clue

It narrows the fault scope to:

prefill writes the cache
→ cache persists
→ decode metadata
→ decode cache reader
→ decode attention

instead of anywhere in the entire model.

16.4 Build a layered chain of numerical evidence

Valid diagnosis sequence for this round:

Q finite?
indices/indptr correct?
raw scale bytes plausible?
decoded K finite?
logits finite?
probs finite?
output token normal?

This is much more efficient than just looking at the final gibberish.

16.5 Record negative results

SwiGLU repair, chat template repair, removal .contiguous() are all valuable:

The full report cannot retain only the last two lines of diff.


17. Conclusion

As of 2026-08-26, vLLM 0.27.x runs on Hygon DCU gfx936, and DeepSeek V4 Flash FP4 completes end-to-end inference through the V1 Engine at TP=4 and TP=8. The work has entered the stability and performance optimization phase.

The ultimate success was achieved not by falling back to a pure BF16 model, but by staying:

The decisive fix was to make the gfx936 Sparse MLA PyTorch decode fallback read the KV cache according to the block-packed physical layout of the C++ writer.

The final evidence also covers:

can therefore be officially confirmed:

vLLM 0.27.x runs on Hygon DCU gfx936. DeepSeek V4 Flash FP4 completed end-to-end testing with the actual weights, the V1 Engine, TP=4/8, and continuous decoding. The next phase is to improve stability at larger batch sizes, validate accuracy, and optimize performance.