MiniMax Music FlashAttention failure
Maestro FlashAttention Failure
Having trouble running MiniMax music on Maestro, it only seems to happen with MiniMax music, not MiniMax H3, for example.
I asked my local agent to write a diagnosis, here's what I got:
Fail log
[MiniMax Music3] Qwen attention backend: flash_attention_2
...
File "C:\pinokio\api\Maestro.git\app\env\lib\site-packages\flash_attn\flash_attn_interface.py", line 96, in _flash_attn_forward
out, softmax_lse, S_dmask, rng_state = flash_attn_gpu.fwd(
RuntimeError: CUDA error: no kernel image is available for execution on the device
...
flash_attn_2_cuda.cp310-win_amd64.pyd!PyInit_flash_attn_2_cuda
Why it suddenly fails
Commit 3f25c48 changed the Windows legacy-runtime FlashAttention package from2.8.2 to 2.7.4 because 2.8.2 could install but fail to load against the
launcher's Python 3.10, PyTorch 2.7.1, and CUDA 12.8 runtime. Update then
force-replaced the installed wheel:
- flash-attn==2.8.2
+ flash-attn==2.7.4.post1
The new binary contains only an sm_89 kernel. This machine's RTX A4500 issm_86, so it cannot execute it. The cached old 2.8.2 binary contains onlysm_100 and sm_120, not sm_86, so the wheel swap did not remove working
RTX A4500 support; it changed which unsupported architecture happened to work.
The failure became visible when Music3 was added. Its backend check only verifies
that flash_attn imports and exposes the expected Python functions. Importing
does not prove that the wheel contains a kernel for the installed GPU. The import
succeeds, Music3 selects flash_attention_2, and the first real CUDA call fails
as shown in the log above.
The actual regression is that the launcher installs an architecture-specific
wheel on incompatible GPUs while Music3 treats a successful import as proof of
runtime compatibility.
Other affected devices
Every Windows NVIDIA GPU routed into Maestro's CUDA 12.8 legacy path receives
the same sm_89-only wheel. The affected GPU names below are expanded from
NVIDIA's current and
legacy CUDA tables; RTX A4500
and RTX A5500 are also included from NVIDIA's
Ampere workstation line card
because the current CUDA table omits those two models.
sm_86 - Ampere
NVIDIA A40; NVIDIA A10; NVIDIA A16; NVIDIA A2; NVIDIA RTX A6000; NVIDIA RTX
A5500; NVIDIA RTX A5000; NVIDIA RTX A4500; NVIDIA RTX A4000; NVIDIA RTX A3000;
NVIDIA RTX A2000; GeForce RTX 3090 Ti; GeForce RTX 3090; GeForce RTX 3080 Ti;
GeForce RTX 3080; GeForce RTX 3070 Ti; GeForce RTX 3070; GeForce RTX 3060 Ti;
GeForce RTX 3060; GeForce RTX 3050 Ti; GeForce RTX 3050.
sm_80 - Ampere
NVIDIA A100; NVIDIA A30.
sm_75 - Turing
NVIDIA T4; QUADRO RTX 8000; QUADRO RTX 6000; QUADRO RTX 5000; QUADRO RTX 4000;
QUADRO RTX 3000; QUADRO T2000; NVIDIA T1200; NVIDIA T1000; NVIDIA T600; NVIDIA
T500; NVIDIA T400; GeForce GTX 1650 Ti; NVIDIA TITAN RTX; GeForce RTX 2080 Ti;
GeForce RTX 2080; GeForce RTX 2070; GeForce RTX 2060.
sm_70 - Volta
NVIDIA V100; Quadro GV100; NVIDIA TITAN V.
sm_61 - Pascal
Tesla P40; Tesla P4; Quadro P6000; Quadro P5200; Quadro P5000; Quadro P4200;
Quadro P4000; Quadro P3200; Quadro P3000; Quadro P2200; Quadro P2000; Quadro
P1000; Quadro P620; Quadro P600; Quadro P500; Quadro P400; P620; P520; NVIDIA
TITAN Xp; NVIDIA TITAN X; GeForce GTX 1080 Ti; GeForce GTX 1080; GeForce GTX
1070 Ti; GeForce GTX 1070; GeForce GTX 1060; GeForce GTX 1050.
sm_60 - Pascal
Tesla P100; Quadro GP100.
sm_52 - Maxwell
Tesla M60; Tesla M40; Quadro M6000 24GB; Quadro M6000; Quadro M5000; Quadro
M4000; Quadro M2000; Quadro M5500M; Quadro M2200; Quadro M620; GeForce GTX
TITAN X; GeForce GTX 980 Ti; GeForce GTX 980; GeForce GTX 970; GeForce GTX 960;
GeForce GTX 950; GeForce GTX 980M; GeForce GTX 970M; GeForce GTX 965M; GeForce
910M.
sm_50 - Maxwell
Quadro K2200; Quadro K1200; Quadro K620; Quadro M1200; Quadro M520; Quadro
M5000M; Quadro M4000M; Quadro M3000M; Quadro M2000M; Quadro M1000M; Quadro
K620M; Quadro M600M; Quadro M500M; NVIDIA NVS 810; GeForce GTX 750 Ti; GeForce
GTX 750; GeForce GTX 960M; GeForce GTX 950M; GeForce 940M; GeForce 930M;
GeForce GTX 850M; GeForce 840M; GeForce 830M.
Cards below sm_50 fail earlier because this PyTorch CUDA runtime does not
contain their targets. Ada (sm_89), Hopper, Blackwell, and Linux take different
launcher paths or packages, so they are not affected by this specific path.
