Maestro FlashAttention Failure
Having trouble running MiniMax music on Maestro, it only seems to happen with MiniMax music, not MiniMax H3, for example.
I asked my local agent to write a diagnosis, here's what I got:
Fail log
[MiniMax Music3] Qwen attention backend: flash_attention_2
...
File "C:\pinokio\api\Maestro.git\app\env\lib\site-packages\flash_attn\flash_attn_interface.py", line 96, in _flash_attn_forward
out, softmax_lse, S_dmask, rng_state = flash_attn_gpu.fwd(
RuntimeError: CUDA error: no kernel image is available for execution on the device
...
flash_attn_2_cuda.cp310-win_amd64.pyd!PyInit_flash_attn_2_cuda
Why it suddenly fails
Commit 3f25c48 changed the Windows legacy-runtime FlashAttention package from
2.8.2 to 2.7.4 because 2.8.2 could install but fail to load against the
launcher's Python 3.10, PyTorch 2.7.1, and CUDA 12.8 runtime. Update then
force-replaced the installed wheel:
- flash-attn==2.8.2
+ flash-attn==2.7.4.post1
The new binary contains only an sm_89 kernel. This machine's RTX A4500 is
sm_86, so it cannot execute it. The cached old 2.8.2 binary contains only
sm_100 and sm_120, not sm_86, so the wheel swap did not remove working
RTX A4500 support; it changed which unsupported architecture happened to work.
The failure became visible when Music3 was added. Its backend check only verifies
that flash_attn imports and exposes the expected Python functions. Importing
does not prove that the wheel contains a kernel for the installed GPU. The import
succeeds, Music3 selects flash_attention_2, and the first real CUDA call fails
as shown in the log above.
The actual regression is that the launcher installs an architecture-specific
wheel on incompatible GPUs while Music3 treats a successful import as proof of
runtime compatibility.
Other affected devices
Every Windows NVIDIA GPU routed into Maestro's CUDA 12.8 legacy path receives
the same sm_89-only wheel. The affected GPU names below are expanded from
NVIDIA's current and
legacy CUDA tables; RTX A4500
and RTX A5500 are also included from NVIDIA's
Ampere workstation line card
because the current CUDA table omits those two models.
sm_86 - Ampere
NVIDIA A40; NVIDIA A10; NVIDIA A16; NVIDIA A2; NVIDIA RTX A6000; NVIDIA RTX
A5500; NVIDIA RTX A5000; NVIDIA RTX A4500; NVIDIA RTX A4000; NVIDIA RTX A3000;
NVIDIA RTX A2000; GeForce RTX 3090 Ti; GeForce RTX 3090; GeForce RTX 3080 Ti;
GeForce RTX 3080; GeForce RTX 3070 Ti; GeForce RTX 3070; GeForce RTX 3060 Ti;
GeForce RTX 3060; GeForce RTX 3050 Ti; GeForce RTX 3050.
sm_80 - Ampere
NVIDIA A100; NVIDIA A30.
sm_75 - Turing
NVIDIA T4; QUADRO RTX 8000; QUADRO RTX 6000; QUADRO RTX 5000; QUADRO RTX 4000;
QUADRO RTX 3000; QUADRO T2000; NVIDIA T1200; NVIDIA T1000; NVIDIA T600; NVIDIA
T500; NVIDIA T400; GeForce GTX 1650 Ti; NVIDIA TITAN RTX; GeForce RTX 2080 Ti;
GeForce RTX 2080; GeForce RTX 2070; GeForce RTX 2060.
sm_70 - Volta
NVIDIA V100; Quadro GV100; NVIDIA TITAN V.
sm_61 - Pascal
Tesla P40; Tesla P4; Quadro P6000; Quadro P5200; Quadro P5000; Quadro P4200;
Quadro P4000; Quadro P3200; Quadro P3000; Quadro P2200; Quadro P2000; Quadro
P1000; Quadro P620; Quadro P600; Quadro P500; Quadro P400; P620; P520; NVIDIA
TITAN Xp; NVIDIA TITAN X; GeForce GTX 1080 Ti; GeForce GTX 1080; GeForce GTX
1070 Ti; GeForce GTX 1070; GeForce GTX 1060; GeForce GTX 1050.
sm_60 - Pascal
Tesla P100; Quadro GP100.
sm_52 - Maxwell
Tesla M60; Tesla M40; Quadro M6000 24GB; Quadro M6000; Quadro M5000; Quadro
M4000; Quadro M2000; Quadro M5500M; Quadro M2200; Quadro M620; GeForce GTX
TITAN X; GeForce GTX 980 Ti; GeForce GTX 980; GeForce GTX 970; GeForce GTX 960;
GeForce GTX 950; GeForce GTX 980M; GeForce GTX 970M; GeForce GTX 965M; GeForce
910M.
sm_50 - Maxwell
Quadro K2200; Quadro K1200; Quadro K620; Quadro M1200; Quadro M520; Quadro
M5000M; Quadro M4000M; Quadro M3000M; Quadro M2000M; Quadro M1000M; Quadro
K620M; Quadro M600M; Quadro M500M; NVIDIA NVS 810; GeForce GTX 750 Ti; GeForce
GTX 750; GeForce GTX 960M; GeForce GTX 950M; GeForce 940M; GeForce 930M;
GeForce GTX 850M; GeForce 840M; GeForce 830M.
Cards below sm_50 fail earlier because this PyTorch CUDA runtime does not
contain their targets. Ada (sm_89), Hopper, Blackwell, and Linux take different
launcher paths or packages, so they are not affected by this specific path.
Maestro FlashAttention Failure
Having trouble running MiniMax music on Maestro, it only seems to happen with MiniMax music, not MiniMax H3, for example.
I asked my local agent to write a diagnosis, here's what I got:
Fail log
Why it suddenly fails
Commit
3f25c48changed the Windows legacy-runtime FlashAttention package from2.8.2to2.7.4because2.8.2could install but fail to load against thelauncher's Python 3.10, PyTorch 2.7.1, and CUDA 12.8 runtime. Update then
force-replaced the installed wheel:
The new binary contains only an
sm_89kernel. This machine's RTX A4500 issm_86, so it cannot execute it. The cached old2.8.2binary contains onlysm_100andsm_120, notsm_86, so the wheel swap did not remove workingRTX A4500 support; it changed which unsupported architecture happened to work.
The failure became visible when Music3 was added. Its backend check only verifies
that
flash_attnimports and exposes the expected Python functions. Importingdoes not prove that the wheel contains a kernel for the installed GPU. The import
succeeds, Music3 selects
flash_attention_2, and the first real CUDA call failsas shown in the log above.
The actual regression is that the launcher installs an architecture-specific
wheel on incompatible GPUs while Music3 treats a successful import as proof of
runtime compatibility.
Other affected devices
Every Windows NVIDIA GPU routed into Maestro's CUDA 12.8 legacy path receives
the same
sm_89-only wheel. The affected GPU names below are expanded fromNVIDIA's current and
legacy CUDA tables; RTX A4500
and RTX A5500 are also included from NVIDIA's
Ampere workstation line card
because the current CUDA table omits those two models.
sm_86- AmpereNVIDIA A40; NVIDIA A10; NVIDIA A16; NVIDIA A2; NVIDIA RTX A6000; NVIDIA RTX
A5500; NVIDIA RTX A5000; NVIDIA RTX A4500; NVIDIA RTX A4000; NVIDIA RTX A3000;
NVIDIA RTX A2000; GeForce RTX 3090 Ti; GeForce RTX 3090; GeForce RTX 3080 Ti;
GeForce RTX 3080; GeForce RTX 3070 Ti; GeForce RTX 3070; GeForce RTX 3060 Ti;
GeForce RTX 3060; GeForce RTX 3050 Ti; GeForce RTX 3050.
sm_80- AmpereNVIDIA A100; NVIDIA A30.
sm_75- TuringNVIDIA T4; QUADRO RTX 8000; QUADRO RTX 6000; QUADRO RTX 5000; QUADRO RTX 4000;
QUADRO RTX 3000; QUADRO T2000; NVIDIA T1200; NVIDIA T1000; NVIDIA T600; NVIDIA
T500; NVIDIA T400; GeForce GTX 1650 Ti; NVIDIA TITAN RTX; GeForce RTX 2080 Ti;
GeForce RTX 2080; GeForce RTX 2070; GeForce RTX 2060.
sm_70- VoltaNVIDIA V100; Quadro GV100; NVIDIA TITAN V.
sm_61- PascalTesla P40; Tesla P4; Quadro P6000; Quadro P5200; Quadro P5000; Quadro P4200;
Quadro P4000; Quadro P3200; Quadro P3000; Quadro P2200; Quadro P2000; Quadro
P1000; Quadro P620; Quadro P600; Quadro P500; Quadro P400; P620; P520; NVIDIA
TITAN Xp; NVIDIA TITAN X; GeForce GTX 1080 Ti; GeForce GTX 1080; GeForce GTX
1070 Ti; GeForce GTX 1070; GeForce GTX 1060; GeForce GTX 1050.
sm_60- PascalTesla P100; Quadro GP100.
sm_52- MaxwellTesla M60; Tesla M40; Quadro M6000 24GB; Quadro M6000; Quadro M5000; Quadro
M4000; Quadro M2000; Quadro M5500M; Quadro M2200; Quadro M620; GeForce GTX
TITAN X; GeForce GTX 980 Ti; GeForce GTX 980; GeForce GTX 970; GeForce GTX 960;
GeForce GTX 950; GeForce GTX 980M; GeForce GTX 970M; GeForce GTX 965M; GeForce
910M.
sm_50- MaxwellQuadro K2200; Quadro K1200; Quadro K620; Quadro M1200; Quadro M520; Quadro
M5000M; Quadro M4000M; Quadro M3000M; Quadro M2000M; Quadro M1000M; Quadro
K620M; Quadro M600M; Quadro M500M; NVIDIA NVS 810; GeForce GTX 750 Ti; GeForce
GTX 750; GeForce GTX 960M; GeForce GTX 950M; GeForce 940M; GeForce 930M;
GeForce GTX 850M; GeForce 840M; GeForce 830M.
Cards below
sm_50fail earlier because this PyTorch CUDA runtime does notcontain their targets. Ada (
sm_89), Hopper, Blackwell, and Linux take differentlauncher paths or packages, so they are not affected by this specific path.