Problem description
Device AddressSanitizer reports a heap-buffer-overflow in the MIOpen CTC loss kernel CTCLossGPU on MI350X (gfx950). The same kernel failure was observed independently in two comprehensive-test shards pinned to different GPUs and using different CTC inputs.
Expected: the CTC tests complete without an out-of-bounds device access.
Actual:
ERROR: AddressSanitizer: heap-buffer-overflow on amdgpu device 0
READ of size 4
#0 CTCLossGPU at ??:0:0
Observed cases:
Smoke/GPU_CTC_FP32.TestFloat/batchSize_32_inputLen_100_labelLen_40_numClass_28_is_softmax_applied_false_blank_id_0_test_id_3
workgroup id (6,0,24)
Smoke/GPU_CTC_FP32.TestFloat/batchSize_1_inputLen_100_labelLen_40_numClass_28_is_softmax_applied_true_blank_id_0_test_id_0
workgroup id (0,0,1)
Environment
- Ubuntu 24.04.4 LTS container; kernel 5.15.0-70-generic
- AMD Instinct MI350X,
gfx950
- Loaded amdgpu version:
7.1.3.31500000
- TheRock ASAN nightly artifact version:
10.2.0
- Public artifact workflow run:
ROCm/rockrel#34546327719
- TheRock artifact commit:
887484afa16ff2200bca0404c3a68321bed8ef9c
- rocm-libraries commit:
512b7a5c1d7eb622c73ea3840a47548144956892
- Test-runner commit:
68c094f7c2481aab4a90bd648fb73bcc936048d1
HSA_XNACK=1
ASAN_OPTIONS=detect_odr_violation=0:quarantine_size_mb=600:symbolize=1:detect_leaks=1
Steps to reproduce
Install the gfx950 ASAN artifacts from the workflow run above, then run the installed MIOpen comprehensive suite with one visible GPU:
export OUTPUT_ARTIFACTS_DIR=<asan-prefix>
export THEROCK_BIN_DIR=<asan-prefix>/bin
export AMDGPU_FAMILIES=gfx950-dcgpu
export AMDGPU_TARGETS=gfx950
export ARTIFACT_GROUP=gfx950-dcgpu
export BUILD_VARIANT=asan-debug
export TEST_TYPE=comprehensive
export TEST_COMPONENT=miopen
export HSA_XNACK=1
export ASAN_SYMBOLIZER_PATH=<asan-prefix>/lib/llvm/bin/llvm-symbolizer
export ASAN_OPTIONS=detect_odr_violation=0:quarantine_size_mb=600:symbolize=1:detect_leaks=1
export LD_PRELOAD=<asan-prefix>/lib/llvm/lib/clang/24/lib/linux/libclang_rt.asan-x86_64.so
ctest -L '^comprehensive$' -L '^ex_gpu_gfx950$' --output-on-failure --parallel 1 --timeout 28800 --test-dir <asan-prefix>/bin/MIOpen -V --tests-information 1,,4
The second observation used --tests-information 2,,4.
Additional evidence
- Shard 1 ran for 3,180 seconds and exited nonzero after the ASAN report.
- The device PC resolves to
CTCLossGPU.
- The two observations used different test inputs and independent GPU processes.
Sanitized full logs can be provided if needed.
Build and test binary locations
Problem description
Device AddressSanitizer reports a heap-buffer-overflow in the MIOpen CTC loss kernel
CTCLossGPUon MI350X (gfx950). The same kernel failure was observed independently in two comprehensive-test shards pinned to different GPUs and using different CTC inputs.Expected: the CTC tests complete without an out-of-bounds device access.
Actual:
Observed cases:
Environment
gfx9507.1.3.3150000010.2.0ROCm/rockrel#34546327719887484afa16ff2200bca0404c3a68321bed8ef9c512b7a5c1d7eb622c73ea3840a4754814495689268c094f7c2481aab4a90bd648fb73bcc936048d1HSA_XNACK=1ASAN_OPTIONS=detect_odr_violation=0:quarantine_size_mb=600:symbolize=1:detect_leaks=1Steps to reproduce
Install the gfx950 ASAN artifacts from the workflow run above, then run the installed MIOpen comprehensive suite with one visible GPU:
The second observation used
--tests-information 2,,4.Additional evidence
CTCLossGPU.Sanitized full logs can be provided if needed.
Build and test binary locations