Porting and Debugging vLLM 0.27.x for DeepSeek V4 Flash FP4 on Hygon DCU gfx936
2026-08-26 Core results:vLLM 0.27.x has been successfully ported to Hygon DCU gfx936. Following targeted DTK adaptations, the V1 Engine completed end-to-end inference with the actual DeepSeek V4 Flash FP4 weights at TP=4 and TP=8. The mixed-precision path—including Sparse MLA, MXFP4 MoE, and continuous decoding—runs end to end. Work has moved from basic functional porting to stability and performance engineering.
0. Supplementary verification results (2026-08-24 to 08-26)
0.1 From functional validation to a stable performance baseline
- The vLLM V1 Engine has been verified; four cards can load the full model.
- Sparse MLA has been pushed from PyTorch reference to gfx936 Triton manual OCP E4M3 decode.
- Full 256-byte decoding reaches
max_abs=0andbad_count=0. The measured single-layer cosine similarity against the real cache is0.99999750, about3.30×the PyTorch reference value. - For M=1, MoE routing Q1 reaches
max_abs=0,cosine=1, andbit_equal=True; it is part of the stable configuration.
0.2 Performance and profile
| Configuration | results | Conclusion |
|---|---|---|
| TP=4, single request, fixed 32 tokens | About 0.99 tok/s | Current recommended baseline |
| TP=8, single request, fixed 32 tokens | About 0.68–0.81 tok/s | Communication overhead offsets additional computing resources |
| TP=4, batch=8 historical measurement | About 2.5 aggregate tok/s | Failed correctness checks; excluded from official performance results |
The clean profile attributes about 30% of steady-state GPU time to the shared MoE forward pass and about 14% to tensor-parallel communication. Many smaller operations (to, copy, sum, add, zeros, and repeat_interleave) also contribute. Disabling routing Q1 increased fixed-generation time from about 67.9 s to 101.3 s, so Q1 must remain enabled. The shared-expert auxiliary-stream change regressed end-to-end performance and was abandoned.
0.3 Batch correctness limits
Batch sizes 1–7 do not show the structural BOS failure. At batch size 8, long decoding can produce a nondeterministic BOS token, degraded output, or NaNs. The issue also occurs with the PyTorch Sparse reader; MXFP4 tests for M=1–9 pass bitwise, and NaNs have been observed before the sampler. The leading areas to investigate are asynchronous streams, workspace reuse, and kernel races—not simply the Triton reader or M=8 GEMM.
The new Q1 cast-elision and grouped-batch MoE paths remain disabled by default. Before either is recommended, it must pass checks against real operator outputs, end-to-end correctness, and a fixed-workload throughput A/B comparison. The latest throughput scan could not run because of a scheduling-account error, so it produced no new performance evidence.
0.4 Stable configuration
DSV4_USE_TRITON_SPARSE=1
DSV4_MOE_ROUTING_Q1=1
DSV4_FP8_WEIGHT_CACHE=0
DSV4_FP8_TRITON_GEMM=0
DSV4_ROCM_SHARED_EXPERTS_STREAM=0
DSV4_GREEDY_LOCAL_ARGMAX=0
DSV4_MOE_Q1_PREPARE=0
VLLM_ROCM_USE_SKINNY_GEMM=0
DSV4_MOE_BACKEND=emulation
Baseline report date: 2026-08-21; latest update: 2026-08-26
Final status: End-to-end inference passed, with coherent output during continuous decoding
Baseline verification job:J17—COMPLETED / ExitCode 0:0
Target platform: eight Hygon DCU gfx936 cards; DTK 26.04
Inference framework: vLLM0.27.2.dev0+g6e448d0ea.d20260819
Model: DeepSeek-V4-Flash-0731
1. Executive Summary
The goal was not merely to load the model weights, but to run DeepSeek V4 Flash FP4 end to end through vLLM on Hygon DCU gfx936:
- Initialize tensor parallelism across eight cards;
- Load approximately 155 GiB of model weights across 48 shards;
- Run FP4 MoE expert weights through the MXFP4 emulation path;
- Dequantize FP8 backbone weights in software on gfx936;
- Use the
fp8_ds_mlaKV cache for DeepSeek V4 Sparse MLA; - Continue decoding after prefill;
- Produce semantically coherent text;
- Complete the Slurm workload with an actual exit code of 0.
Output from the successful run:
TOKEN_IDS: [223, 30594, 9790, 303, 70037, 12203, 10706, 804, 1175,
10539, 637, 53897, 96240, 2996, 84998, 666, 3118, 637,
58368, 621, 666, phase, 53091, 4374, 1465, 223, 1959,
17585, 42108, 303, 56463, 10162]
TOKEN_DECODE:
[' ', '你好', '呀', ' ,', '很高兴', '在这里', '遇到', '你', '!',
'让我', '来', '做个', '自我介绍', '吧', '~\n\n', '**', '初', '来',
'乍', '到', '**', ':', 'Deep', 'Se', 'ek', ' ', '这个', '版本',
'的我', ' ,', '确实是', '有点']
OUTPUT_REPR:
' 你好呀 ,很高兴在这里遇到你!让我来做个自我介绍吧~\n\n
**初来乍到**:DeepSeek 这个版本的我 ,确实是有点'
DSV4_VERDICT: COHERENT
Final job status:
JobID|State|ExitCode|Elapsed|MaxRSS
J17|COMPLETED|0:0|00:05:15|
J17.batch|COMPLETED|0:0|00:05:15|65899600K
J17.extern|COMPLETED|0:0|00:05:15|0
The root cause was not the FP4 MoE numerical format, chat template, skinny GEMM, or tokenizer. It was the gfx936 Sparse MLA PyTorch decode fallback misreading the fp8_ds_mla KV cache.
KV cache is:
[num_blocks, block_size, 584]
But each block is not 64 independent 584-byte tokens arranged in sequence in physical memory, but:
[64 × 576-byte token data][64 × 8-byte UE8M0 scales]
The original reader incorrectly assumed that each token occupied 584 contiguous bytes. As a result, it:
- treats the data of other tokens as scale;
- interpreted scale bytes as BF16 RoPE values;
- produced NaNs and values near
1e37; - attention logits and softmax probabilities become NaN;
- The first token is generated by prefill and may appear normal;
- The second token enters a corrupted decode state, after which generation repeatedly selects token 0 (BOS).
After the fix, the reader and C++ writer use matching address calculations:
data = block_base + token_pos * 576
scale = block_base + block_size * 576 + token_pos * 8
Verification shows:
scale_bytes = [118, 119, 119, 119, 118, 119, 118, 0]
NoPE finite = True
RoPE finite = True
attention logits finite = True
attention probs finite = True
Together, these checks form a complete evidence chain from physical addresses and quantization scales through KV-cache values, attention logits, and token IDs to the final natural-language output.
2. Goals, scope, and success criteria
2.1 Core Objectives
The core goal has always been:
runs through DeepSeek V4 Flash FP4 inference through vLLM on Hygon DCU gfx936.
The "FP4" here specifically refers to the model's MoE expert using MXFP4 weights. Not all tensors of the model are FP4; the actual checkpoint also contains:
- MXFP4 expert weight;
- FP8 E4M3 backbone weight;
- UE8M0 scale;
- BF16 activation, RoPE and some intermediate states.
Therefore, the real goal is to run through the mixed precision execution chain rather than verifying an FP4 GEMM alone.
2.2 Outcomes that do not constitute completion
The following states do not count as a successful end-to-end run:
- can only import vLLM;
- Only
DeepseekV4ForCausalLM; - can only read model config;
- can only load all weights;
- Show only
ENGINE_READY; - only generates the first token;
- output is empty;
- continuously generates BOS, repeated tokens or garbled characters;
- Python fails but Slurm shows
COMPLETED; - only verifies BF16 and does not execute the FP4 expert path.
2.3 Final success criteria
The end-to-end success criteria used in this round are:
- [x] 8 gfx936 are visible;
- [x] TP=8 Initialization successful;
- [x] Complete weight loading successful;
- [x] expert dtype confirmed to be FP4;
- [x] MXFP4 emulation backend initialization successful;
- [x] FP8 software inverse quantization path work;
- [x]
fp8_ds_mlacache initialization successful; - [x] prefill completed;
- [x] Perform real decode at least once;
- [x] decode attention of Q/K/logits/probs fully finite;
- [x] generates 32 non-degraded tokens;
- [x] The output text is semantically coherent;
- [x] Workload exit code is 0;
- [x] Slurm main job and batch step are both
COMPLETED.
3. Hardware and software environment
3.1 Computing resources
Measured environment:
| Project | value |
|---|---|
| accelerator | Hygon DCU |
| Architecture | gfx936:sramecc+:xnack- |
| Number of cards | 8 |
| GPU memory per card (as reported by PyTorch) | Approximately 68.7 GB decimal, which is approximately 63.98 GiB |
| Slurm Partition | 内部 DCU 分区 |
| CPU | --cpus-per-task=64 |
| working memory | --mem=240gb |
| Maximum test time | 02:00:00 |
3.2 Software Stack
| Component | version or path |
|---|---|
| DTK | 26.04 |
| GCC module | 12.2.0 |
| Python | 3.12 ,vLLM venv |
| vLLM | 0.27.2.dev0+g6e448d0ea.d20260819 |
| model | DeepSeek-V4-Flash-0731 |
| vLLM source code | <VLLM_SOURCE> |
| venv | <VLLM_ENV> |
| Test Catalog | <TEST_WORKSPACE> |
3.3 Model and quantization configuration
Read from the running log:
quant_config:
{
'activation_scheme': 'dynamic',
'fmt': 'e4m3',
'quant_method': 'fp8',
'scale_fmt': 'ue8m0',
'weight_block_size': [128, 128]
}
vLLM runtime confirmation:
DeepSeek V4 expert_dtype resolved to 'fp4'
Using 'EMULATION' Mxfp4 MoE backend.
Using DeepSeek's fp8_ds_mla KV cache format.
Using FP8 indexer cache for Lightning Indexer.
3.4 Runtime environment variables
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export TORCHINDUCTOR_USE_STATIC_CUDA_LAUNCHER=0
export DSV4_SHIM_DEBUG=1
export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1
export DSV4_TP=8
export DSV4_MAXLEN=512
export DSV4_MEMUTIL=0.90
export DSV4_KV_CACHE_DTYPE=fp8
export DSV4_MOE_BACKEND=emulation
export VLLM_ROCM_USE_SKINNY_GEMM=0
Also used during debugging:
export DSV4_SPARSE_DEBUG=1
4. Implementation approach
The end-to-end execution chain can be simplified to:
Model config / tokenizer
│
▼
FP8 backbone weights ──software dequantization──► BF16 GEMM
│
├──► Q / KV / indexer / compressor
│
▼
FP4 MoE expert ──MXFP4 emulation──► MoE output
│
▼
DeepSeek V4 Sparse MLA
│
├── prefill
├── SWA cache write(C++ fused writer)
├── compressed/top-k cache
└── decode fallback(gfx936 PyTorch reference)
│
▼
logits → softmax → attention output
│
▼
LM head → next token
The investigation uncovered four independent compatibility and correctness issues:
- Triton / DAS Triton API and lowering compatibility issues;
- A device-side assertion in gfx936 skinny GEMM;
- Input format misjudgment caused by prompt/chat template;
- The physical layout of the Sparse MLA KV cache reader is wrong.
Category 4 is the final root cause of "the first token is normal, and subsequent tokens are all BOS".
5. Diagnostic methods and test probes
The diagnostic probe reports progress by stage:
### STAGE: env
### STAGE: config
### STAGE: vllm import + registry
### STAGE: engine init (the real test)
### STAGE: generate
Answer each stage separately:
| Stage | Questions to answer |
|---|---|
| env | Can PyTorch detect all eight gfx936 devices, and are their memory sizes reported correctly? |
| config | Can the model architecture, layer count, expert configuration, and quantization settings be read correctly? |
| registry | Is DeepSeek V4 registered with vLLM? |
| engine init | Did weight loading, tensor parallelism, MoE setup, KV-cache allocation, and warm-up complete? |
| generate | Do prefill and continuous decoding produce valid tokens? |
Print during generation:
TOKEN_IDS
FINISH_REASON
TOKEN_DECODE
OUTPUT_REPR
OUTPUT
DSV4_VERDICT
Slurm script explicitly saves Python return code:
python load_dsv4.py > "$RAW" 2>&1
rc=$?
echo "PY_EXIT=$rc"
exit "$rc"
So the Slurm status, Python verdict and raw logs in the report can be cross-validated.
6. Validation timeline and job records
6.1 Summary table
| Job | Status | Time | reaches level | Main results |
|---|---|---|---|---|
| J01 | FAILED 1 | 06:18 | engine init | Triton lowering PassManager::run failed |
| J02 | FAILED 1 | 06:18 | engine init | Triton softmax(dim=...) API is not compatible with |
| J03 | FAILED 1 | 06:41 | engine init | Similar softmax lowering problem |
| J04 | FAILED 1 | 06:16 | engine init | Similar softmax lowering problem |
| J05 | FAILED 1 | 06:45 | engine init | fallback still triggers softmax API issue |
| J06 | FAILED 1 | 06:43 | generate | ENGINE_READY;skinny GEMM device assertion |
| J07 | FAILED 2 | 06:49 | generate | Disable skinny GEMM; no hard crashes, output degradation |
| J08 | COMPLETED 0 | 07:11 | generate | [32,0,0,...]; Probe false alarm COHERENT |
| J09 | FAILED 2 | 03:32 | generate | Still [0,0,...] |
| J10 | CANCELLED | 00:16 | Start | The old script submitted by mistake, proactively canceled |
| J11 | FAILED 2 | 07:13 | generate | chat template repair; first token 你好, follow-up BOS |
| J12 | FAILED 1 | 01:15 | engine init | auto KV dtype was replaced by fp8_ds_mla layout reject |
| J13 | FAILED 2 | 03:05 | generate | confirms again that the first token is random and subsequent BOS |
| J14 | FAILED 2 | 03:06 | generate | diagnostics are consumed by warmup and the real decode |
| J15 | FAILED 2 | 04:51 | generate | captures the real decode: cache K/attention full NaN |
| J16 | FAILED 2 | 07:58 | generate | The first stride repair is insufficient and still NaN |
| J17 | COMPLETED 0 | 05:15 | end-to-end | The physical layout was corrected, and the output is consistent |
Interpretation:
FAILED 2usually means the probe correctly classified degraded output as a failure; it does not indicate a hardware crash.J08returnedCOMPLETED 0despite clearly incorrect output, exposing a flaw in the probe's early success check.- The recurring
libcudart is not loadedmessage is ROCm cleanup noise, not the first error that caused the run to fail.
6.2 J01: Triton lowering blocks engine startup
RuntimeError: PassManager::run failed
RuntimeError: Engine core initialization failed.
At this stage:
- The model, PyTorch, and worker have started.
- The failure is in compilation/lowering, not in model-file availability.
- Unsupported Triton implementations on gfx936/DAS must be replaced individually.
6.3 J02–J05: Resolve softmax API incompatibility
Stable reproduction of multi-round operations:
triton.compiler.errors.CompilationError
TypeError("softmax() got an unexpected keyword argument 'dim'")
This is the difference between the upstream code and the current DAS Triton API. At this stage, by adjusting fallback and avoiding incompatible Triton paths, the failure position is gradually advanced to the real engine warmup.
Engineering lesson: Avoid repeatedly patching isolated Triton syntax such as tt.dot_scaled. For operations that cannot be lowered on gfx936, provide a testable PyTorch/BF16 reference fallback.
6.4 J06: Engine startup succeeds; first device-side error
Key developments:
Model loading took 19.99 GiB memory and 205.150379 seconds
GPU KV cache size: 85,259 tokens
ENGINE_READY in 335.4s
The first device-side error during generation:
csrc/rocm/skinny_gemms.hip:580
wvSplitK_hf_sml_
Device-side assertion `false' failed.
This is a true device-side error followed by a VMFault.
Workaround:
export VLLM_ROCM_USE_SKINNY_GEMM=0
With this setting:
- device assertion disappears;
- VMFault disappears;
- generation completes;
- the failure changes from a hardware crash to incorrect numerical output.
6.5 J07–J08: Crashes resolved, but output remains incorrect
J08:
TOKEN_IDS: [32, 0, 0, ..., 0]
TOKEN_DECODE:
['>', '<|begin▁of▁sentence|>', ...]
OUTPUT_REPR: '>'
The probe initially reported:
DSV4_VERDICT: COHERENT
The probe's success check only tested whether:
- the output was empty; or
- a few characters were repeated conspicuously.
Because OUTPUT_REPR contained only one >, it passed the weak “nonempty and not repetitive” check, producing a false positive.
The acceptance check was strengthened to inspect all of the following:
- token IDs;
- the decoded text for each token;
- the generation finish reason;
- the generated text itself;
- whether BOS repeats in subsequent tokens.
6.6 J09: Add a SwiGLU fallback
The original gfx936 emulation fallback was:
act = F.silu(gate * 1.702) * up
This did not match the model configuration or the vLLM reference. It was changed to:
gate = torch.clamp(gate, max=gemm1_limit)
up = torch.clamp(up, min=-gemm1_limit, max=gemm1_limit)
act = F.silu(gate * gemm1_alpha) * up
and uses:
gemm1_alpha = 1.0
gemm1_limit = 10.0
Job result:
TOKEN_IDS: [0, 0, ..., 0]
Conclusion:
- The SwiGLU fix is required to match the model's activation formula.
- It does not, by itself, resolve the repeated-BOS output.
- Further investigation therefore focused on the prefill/decode boundary.
6.7 J11: Fix the chat template and isolate the decode issue
A plain prompt:
llm.generate(["你好 ,请介绍一下你自己 ."])
does not necessarily follow DeepSeek V4's custom conversation template. After switching to the chat API:
TOKEN_IDS: [30594, 0, 0, ..., 0]
TOKEN_DECODE:
['你好', '<|begin▁of▁sentence|>', ...]
OUTPUT_REPR: '你好'
This is very critical boundary evidence:
- prefill can produce the expected first token (the expected localized greeting);
- the second token is already replaced by BOS;
- the tokenizer and chat template affect the first token;
- independent numerical errors remain during continuous decoding.
6.8 J12: Test BF16 and automatic KV-cache settings
The automatic KV-cache setting was tested:
DSV4_KV_CACHE_DTYPE=auto
Engine initialization rejected it explicitly:
AssertionError:
DeepseekV4 fp8_ds_mla layout only supports fp8 kv-cache, got auto
The current DeepSeek V4 ROCm Sparse MLA implementation is bound to the fp8_ds_mla layout. Setting the cache type to auto cannot bypass this constraint; the layout's read and write paths must be implemented correctly.
6.9 J13: Confirm repeatability
ENGINE_READY in 93.2s
TOKEN_IDS: [57297, 0, 0, ..., 0]
TOKEN_DECODE:
[' Modified', BOS, BOS, ...]
The first generated token varied between runs:
>你好Modified看来- Space
0
But the common pattern is always:
[prefill token, BOS, BOS, BOS, ...]
Therefore the problem cannot be interpreted as a fixed tokenizer mapping; the failure occurs in the decode calculation chain.
6.10 J14: Initial diagnostic probe misses the failing decode
First diagnostic output:
q shape = (8, 8, 512)
main_indptr = [0, 0, ..., 0]
This is the dummy invocation of engine warmup, not a real request. The one-time diagnostic flag is consumed by warmup, causing the real decode to not be printed.
Corrected trigger condition:
diag_once = (
diag
and num_queries <= 2
and not getattr(fn, "DSV4_DIAG_ONCE", False)
)
num_queries=1, warmup's num_queries=8.
6.11 J15: Numerical evidence identifying the root cause
Metadata from the failing decode:
q shape = (1, 8, 512)
main_cache shape = (35644, 64, 584)
main_cache stride= (1002240, 584, 1)
main_indices = [128, 129, ..., 138, ...]
main_indptr = [0, 11]
scale = 0.04419417382415922
qfinite = True
This proves:
- Q is finite;
- The number of indexes is 11, which is consistent with the prompt context length;
- query itself is not a source of NaN.
After decoding, however, the cache contained:
scale_bytes:
[87, 245, 226, 232, 107, 107, 156, 110]
nope_finite = False
nope_max = nan
rope_finite = True
rope_max = 1.0135363467859984e+37
Attention diagnostics:
logits_finite = False
logits_min = nan
logits_max = nan
probs_finite = False
probs_max = nan
The complete failure chain is:
Q finite
→ cache decode error
→ NoPE NaN / RoPE 1e37
→ attention logits NaN
→ softmax NaN
→ hidden-state corruption
→ LM-head degradation
→ token 0 / BOS
6.12 J16: Why the first stride fix was insufficient
The initial fix made the cache contiguous:
cache.view(torch.uint8).contiguous()
That change was removed. The next attempt kept the original cache view and read it with:
raw[block, pos]
The result remained incorrect:
scale_bytes = [88, 245, 226, 232, ...]
nope_finite = False
rope_max ≈ 1e37
This showed that the issue was not simply a lost stride(0) caused by .contiguous().
Comparison with the C++ writer showed that stride(1)=584 cannot be used by itself to locate individual tokens. The physical cache is block-packed rather than token-interleaved.
6.13 J17: Correct the physical cache layout
Final diagnosis:
scale_bytes:
[118, 119, 119, 119, 118, 119, 118, 0]
nope_finite = True
nope_max = 1.875
rope_finite = True
rope_max = 5.75 ~ 5.78125
Attention values were valid across tensor-parallel ranks. For example:
TP0:
logits_finite = True
logits_min = -2.2683308124542236
logits_max = 6.992861270904541
probs_finite = True
probs_max = 0.7818039655685425
TP2:
logits_finite = True
logits_min = -3.1048312187194824
logits_max = 4.642844200134277
probs_finite = True
probs_max = 0.5944044589996338
TP7:
logits_finite = True
logits_min = -2.159546136856079
logits_max = 6.589494228363037
probs_finite = True
probs_max = 0.9574916362phase
The final token no longer degrades and forms coherent Chinese.
7. Root cause: logical shape and physical KV-cache layout diverge
7.1 Logical tensor representation
The backend declares that each token occupies 584 bytes:
448 byte FP8 NoPE
128 byte BF16 RoPE
8 byte UE8M0 scales/padding
logical shape:
[num_blocks, block_size, 584]
This is easy for the reader to write:
token = cache[block, pos]
data = token[:576]
scale = token[576:584]
But this explanation is wrong.
7.2 Addresses written by the C++ writer
writer uses:
block_base =
k_cache + block_idx * kv_block_stride;
token_fp8_ptr =
block_base + pos_in_block * kTokenDataBytes;
token_bf16_ptr =
token_fp8_ptr + kNopeDim;
token_scale_ptr =
block_base
+ cache_block_size * kTokenDataBytes
+ pos_in_block * kScaleBytesPerToken;
Constant:
kNopeDim = 448
kRopeDim = 64 BF16 = 128 byte
kTokenDataBytes = 576
kScaleBytesPerToken = 8
block_size = 64
Therefore the single block physical layout:
offset 0
│
├── token 0 data: 576 B
├── token 1 data: 576 B
├── ...
├── token 63 data: 576 B
│
├── token 0 scale: 8 B
├── token 1 scale: 8 B
├── ...
└── token 63 scale: 8 B
Total valid data:
64 × 576 + 64 × 8
= 36,864 + 512
= 37,376 byte
while running observed:
main_cache.stride() = (1002240, 584, 1)
stride(0) is determined by the allocator/shared cache layout, which is much larger than the valid data of a single block. The reader must retain the real block stride; it cannot first .contiguous().
7.3 Two errors in the original reader
Old code:
cache_u8 = cache.view(torch.uint8).contiguous()
raw = cache_u8.reshape(cache.shape[0], block_size, -1)
token = raw[block, pos]
data = token[:, :576]
scale_bytes = token[:, 576:584]
Error 1:.contiguous() destroys physics stride(0).
Error 2: Even if stride is retained,token[:, 576:584] still assumes that scale follows the data of each token; the actual scale is concentrated at the end of the block.
7.4 Final reader
The final implementation first constructs a byte view that retains block stride:
raw = cache_u8.as_strided(
(cache.shape[0], block_size * 584),
(cache_u8.stride(0), 1),
)
Read data:
data_offsets = (
pos[:, None] * 576
+ torch.arange(576, device=device)[None, :]
)
data = raw[block[:, None], data_offsets]
Reading scale:
scale_offsets = (
block_size * 576
+ pos[:, None] * 8
+ torch.arange(8, device=device)[None, :]
)
scale_bytes = raw[block[:, None], scale_offsets]
UE8M0 decoding:
k_scales = torch.exp2(scale_bytes.float() - 127.0)
FP8 NoPE:
k_nope = data[:, :448].contiguous().view(fp8_dtype).float()
k_nope *= k_scales[:, :7].repeat_interleave(64, dim=1)
BF16 RoPE:
k_rope = (
data[:, 448:576]
.contiguous()
.view(torch.bfloat16)
.float()
)
7.5 Basis for the final scale
Final scale bytes:
[118, 119, 119, 119, 118, 119, 118, 0]
correspond to 7 64-element groups of 448 NoPE elements.
For example:
encoded 118 → scale = 2^(118 - 127) = 2^-9
encoded 119 → scale = 2^(119 - 127) = 2^-8
is padding, the value is 0, and does not participate in the dequantization of the 7 actual groups.
This is exactly the same as the writer’s 7 UE8M0 scale + 1 byte pad design.
8. Hypotheses tested and their limits
8.1 Hypothesis: the FP4 MoE implementation is faulty
The evidence does not support this.
- FP4 backend normal initialization;
- Weight loading completed;
- FP4 MoE is not modified after repairing the cache reader, and the output is restored to consistency;
- Therefore the final root cause of consecutive BOS is not in the FP4 expert main path.
But the SwiGLU parameter fix should still be retained as it fixes model formula consistency.
8.2 Hypothesis: the tokenizer or chat template is the sole cause
is only partially true.
- chat template After repair, the first token changed from a meaningless symbol to "Hello";
- But after the second token it is still all BOS;
- Therefore, template explains the input format problem and cannot explain decode numerical degradation.
8.3 Hypothesis: switching the KV cache to auto/BF16 avoids the issue
is not established.
Current DeepSeek V4 ROCm Sparse MLA layout mandatory requirement:
fp8_ds_mla
auto was explicitly rejected in engine init.
8.4 Hypothesis: removing contiguous() is sufficient
is not established.
.contiguous() does lose block stride, but the logical token stride also cannot represent the physical layout. Must also be fixed:
- block stride ;
- data/scale partition address.
8.5 Hypothesis: a correct first token proves overall numerical correctness
is not established.
The first generated token is produced during prefill; single-token decoding begins only when the next token is requested. The fault occurs at the prefill-to-decode transition.
Therefore end-to-end testing requires at least:
max_tokens >= 2
A safer bet would be 16 or 32.
8.6 Hypothesis: a COMPLETED Slurm job proves model correctness
is not established.
J08 is a typical counterexample:
State = COMPLETED
ExitCode = 0
TOKEN_IDS = [32, 0, 0, ...]
framework does not throw an exception, which does not mean that the value is correct. Compatibility report must also be checked:
- scheduler state ;
- workload exit code ;
- raw traceback ;
- token IDs ;
- output text;
- intermediate numerical finite state.
9. Implemented changes
9.1 gfx936 compatibility shim
sitecustomize.py is automatically applied after worker spawn to make up for the current Triton/vLLM API differences and avoid premature import of Triton at startup.
include:
- Triton language API placeholder/backport ;
reduce_or;- JITFunction metadata ;
- HCU compiler arch default value;
- knobs runtime hooks ;
- vLLM FP8 linear software inverse quantization.
These are the infrastructure that the engine can use to start and complete weight loading.
9.2 Software dequantization for the FP8 backbone
gfx936 does not support native FP8 as expected by vLLM _scaled_mm path, so FP8 linear uses BF16 software inverse quantization fallback.
This path sacrifices performance but establishes a baseline for correctness.
9.3 MXFP4 MoE emulation
DSV4_MOE_BACKEND=emulation
weights remain in FP4 checkpoint format and are dequantized and calculated on demand by the emulation expert path. This path is used to first prove that the model functions correctly.
9.4 SwiGLU parameter alignment
Fixed the error 1.702 formula with the one consistent with the model configuration:
alpha = 1.0
limit = 10.0
9.5 Disable gfx936 skinny GEMM
VLLM_ROCM_USE_SKINNY_GEMM=0
Avoid wvSplitK_hf_sml_ .
9.6 Sparse MLA decode fallback
gfx936 cannot stabilize lower, so PyTorch reference decode is used.
The final critical fix is located at:
vllm/v1/attention/ops/rocm_aiter_mla_sparse.py
The repair content is fp8_ds_mla block-packed cache.
10. Verification results
10.1 Engine initialization
ENGINE_READY in 228.9s
Across different nodes, cache conditions, and JIT states, initialization previously ranged from about 93 to 338 seconds. A single engine-ready measurement therefore cannot establish steady-state performance.
10.2 KV-cache values
All eight tensor-parallel workers reported the same scale bytes:
[118, 119, 119, 119, 118, 119, 118, 0]
and:
NoPE finite = True
NoPE abs max = 1.875
RoPE finite = True
RoPE abs max ≈ 5.75–5.78125
10.3 Attention values
Across all tensor-parallel ranks:
logits_finite = True
probs_finite = True
Logits and maximum softmax probabilities differ slightly across ranks, as expected because each rank holds different attention-head shards.
10.4 Output quality
The 32-token response contains no repeated BOS tokens. Its Chinese text is coherent and includes natural punctuation and Markdown formatting:
你好呀 ,很高兴在这里遇到你!让我来做个自我介绍吧~
**初来乍到**:DeepSeek 这个版本的我 ,确实是有点
10.5 Resources
MaxRSS = 65899600K
This is the batch step's CPU resident memory. The GPU-side log reports:
Model loading took approximately 19.99 GiB per worker/device
GPU KV cache size approximately 85K tokens
These values depend on the node, JIT state, allocator, and cache configuration. They document this functional-validation run and should not be treated as general performance figures.
11. Log analysis: distinguish noise from actionable errors
11.1 Usage-report JSON error
The following message appeared in several logs:
Exception in thread Thread-1 (_report_usage_worker)
json.decoder.JSONDecodeError
The warning originates in the usage-report background thread, not in engine startup or model computation. Similar noise also appears in runs that complete successfully.
11.2 libcudart assertion during ROCm shutdown
Failed jobs commonly emit this message during cleanup:
AssertionError:
libcudart is not loaded in the current process,
try setting VLLM_CUDART_SO_PATH
This is a ROCm compatibility issue with the vLLM shutdown path trying to import the CUDA allocator. It appears after generating results and probe verdicts, and should not be misjudged as the root cause of this inference failure.
11.3 How to read FIRST BLOCKER
Identify the first blocker in this order:
- Find the earliest workload traceback in the raw log;
- Confirm whether it occurs in the main execution path;
- distinguish background-thread and shutdown messages from errors in the main inference path;
- compare
ENGINE_READY,TOKEN_IDS, and the probe verdict; - Finally check the Slurm state and exit code.
12. Reproduction experiment
12.1 Slurm resources
#SBATCH -p <DCU_PARTITION>
#SBATCH --gres=dcu:8
#SBATCH --cpus-per-task=64
#SBATCH --mem=240gb
#SBATCH --time=02:00:00
12.2 Environment
module purge
module load compiler/gcc/12.2.0 compiler/dtk/26.04
source <VLLM_ENV>/bin/activate
export LD_LIBRARY_PATH=<DTK_ROOT>/gcvm/lib:\
<DTK_ROOT>/lib:\
<DTK_ROOT>/lib64:\
<DTK_ROOT>/comgr/lib64:\
$LD_LIBRARY_PATH
export PYTHONPATH=<TEST_WORKSPACE>:\
<TRITON_KERNELS>:\
$PYTHONPATH
12.3 vLLM parameters
llm = LLM(
model=MODEL,
tensor_parallel_size=8,
max_model_len=512,
kv_cache_dtype="fp8",
moe_backend="emulation",
enforce_eager=True,
gpu_memory_utilization=0.90,
trust_remote_code=True,
)
Dialog input should use the model chat template:
llm.chat(
[[
{
"role": "user",
"content": "你好 ,请介绍一下你自己 .",
}
]],
SamplingParams(
max_tokens=32,
temperature=0,
ignore_eos=True,
),
)
12.4 Result checks
Must check:
ENGINE_READY
TOKEN_IDS
TOKEN_DECODE
OUTPUT_REPR
DSV4_VERDICT
PY_EXIT
sacct State / ExitCode
should not just grep OUTPUT:because special tokens such as BOS may be omitted from the final text by the tokenizer.
13. Code version and supporting files
Local warehouse:
| file | Purpose |
|---|---|
probes/load_dsv4.py |
Phased end-to-end loading and generation of probes |
probes/load_dsv4.slurm |
Basic Slurm Script |
probes/sitecustomize.py |
spawn worker automatically compatible with shim |
probes/patch_sparse_diag_min.py |
Sparse MLA Q/cache/attention numerical diagnosis |
probes/patch_sparse_physical_layout.py |
Final physical cache layout repair script |
dsv4-vllm-handoff.md |
Earlier torch/vLLM/mmac porting history |
dsv4-dcu-record.html |
Project process record |
Remote key source code:
<VLLM_SOURCE>/
vllm/v1/attention/ops/rocm_aiter_mla_sparse.py
The remote end retains pre-repair backup and diagnostic scripts for easy comparison.
14. Findings and scope
14.1 Proven
- DeepSeek V4 Flash checkpoint can be fully loaded on a single node 8 × gfx936;
- vLLM TP=8 can be initialized;
- FP4 expert weighting can be performed through MXFP4 emulation;
- FP8 backbone can be executed through software inverse quantization;
- Sparse MLA prefill and decode executable;
fp8_ds_mlacache value is correct after repairing the physical reader;- single request greedy generates 32 tokens and can output coherent Chinese;
- Slurm workload completed normally.
14.2 Not yet proven
- Long context to 512 token full coverage stability;
- Multiple concurrent requests;
- Complex reuse scenario of prefix caching;
- All boundaries of chunked prefill;
- All length combinations of C4/C128 compressed extra cache;
- Multiple rounds of dialogue;
- EOS normal stop behavior;
- temperature/top-p non-greedy sampling;
- long-term service stability;
- GPU memory fragmentation and repeated allocation/reclamation requests;
- The systematic error of mathematical accuracy relative to the official CUDA/BF16 reference;
- production throughput and P50/P95 latency;
- End-to-end performance after replacing emulation with native gfx936 FP4 kernel.
Therefore the current expression should be:
FP4 functional link and continuous decode have been run through and have a correctness baseline; production-level stability, accuracy and performance acceptance still require subsequent testing.
15. Recommended follow-up work
15.1 Priority 1: Make regression tests permanent
Adds at least the following regressions:
- single request, 32 tokens, Chinese chat;
- single request, 128 tokens;
- Two rounds of dialogue;
- 8 different prompts;
- Two concurrent requests;
- prompt length crosses block boundary;
- SWA window border;
- C4/C128 compressor boundary;
- prefix caching switch A/B;
- runs 50 times to check for occasional NaNs.
Must check every time:
all token IDs
all special tokens
finite checks
Slurm exit code
VMFault/device assertion
15.2 Priority 2: Turn diagnostics into test assertions
Upgrade one-time printing to optional assertion:
assert torch.isfinite(k_nope).all()
assert torch.isfinite(k_rope).all()
assert torch.isfinite(logits).all()
assert torch.isfinite(probs).all()
Only in DSV4_SPARSE_VALIDATE=1 is used to avoid loss of production performance.
15.3 Priority 3: Compare numerical accuracy
Recommended layering:
- cache writer/reader round-trip ;
- single token attention output;
- single layer hidden state cosine;
- Hidden state of each layer;
- next-token logits top-k ;
- greedy token sequence ;
- small data set perplexity.
The reference baseline can be:
- CUDA official/upstream implementation;
- BF16 cache/reference attention ;
- CPU small size reference.
15.4 Priority 4: Performance optimization
The current PyTorch decode fallback will produce:
- Advanced Index;
- temporary tensor;
- FP8→FP32 ;
- Python query loop;
- multiple kernel launches;
- host sync type
.item().
It is suitable as a correctness reference, not suitable as a final high-performance implementation.
The next step should be to port the correct physical address formula to:
- gfx936 lowerable Triton kernel; or
- HIP/C++ dedicated decode kernel.
must retain this round of reference implementation as an oracle before optimization.
15.5 Priority 5: Remove temporary debugging code
Before production:
- will
DSV4_SPARSE_DEBUGLeave as default off; - Delete unnecessary synchronous printing;
- preserves physical layout annotations;
- Add unit test coverage block-packed layout;
- Organize patches into reviewable git commits;
- will
VLLM_ROCM_USE_SKINNY_GEMM=0becomes the gfx936 platform default or explicit capability gate.
16. Engineering lessons
16.1 Tensor shape does not determine physical layout
Quantized caches often use logical shapes that "conveniently express capacity", but writers may use completely different physical layouts for vectorization, TMA, coalescing, or scale partitions.
The reader must use the writer's address formula as the source of truth.
16.2 Instrument prefill and decode separately
warmup, prefill and decode may be used differently:
- shape ;
- kernel ;
- cache slot ;
- backend ;
- tensor stride.
The "first function call" is usually just a warmup and does not represent a real request.
16.3 A correct first token is a useful diagnostic clue
It narrows the fault scope to:
prefill writes the cache
→ cache persists
→ decode metadata
→ decode cache reader
→ decode attention
instead of anywhere in the entire model.
16.4 Build a layered chain of numerical evidence
Valid diagnosis sequence for this round:
Q finite?
indices/indptr correct?
raw scale bytes plausible?
decoded K finite?
logits finite?
probs finite?
output token normal?
This is much more efficient than just looking at the final gibberish.
16.5 Record negative results
SwiGLU repair, chat template repair, removal .contiguous() are all valuable:
- Some independent bugs have been fixed;
- Some have narrowed the scope;
- Some proof assumptions are insufficient.
The full report cannot retain only the last two lines of diff.
17. Conclusion
As of 2026-08-26, vLLM 0.27.x runs on Hygon DCU gfx936, and DeepSeek V4 Flash FP4 completes end-to-end inference through the V1 Engine at TP=4 and TP=8. The work has entered the stability and performance optimization phase.
The ultimate success was achieved not by falling back to a pure BF16 model, but by staying:
- FP4 expert ;
- MXFP4 emulation ;
- FP8 backbone software inverse quantization;
- FP8
fp8_ds_mlaKV cache ; - Sparse MLA ;
- TP=8.
The decisive fix was to make the gfx936 Sparse MLA PyTorch decode fallback read the KV cache according to the block-packed physical layout of the C++ writer.
The final evidence also covers:
- scheduler ;
- engine ;
- weight and backend;
- cache bytes ;
- dequantized K ;
- attention logits/probs ;
- token IDs ;
- natural language output;
- workload exit code.
can therefore be officially confirmed:
vLLM 0.27.x runs on Hygon DCU gfx936. DeepSeek V4 Flash FP4 completed end-to-end testing with the actual weights, the V1 Engine, TP=4/8, and continuous decoding. The next phase is to improve stability at larger batch sizes, validate accuracy, and optimize performance.