Pinokio
FeedApps
Download Pinokio
Sign inCreate accountJoin
Sign inCreate accountJoin

Post

Conversation
@designer_99
2h ago

Microsoft Windows [Version 10.0.26200.9457]
(c) Microsoft Corporation. All rights reserved.

C:\pinokio\api\wan.git\app>conda_hook & conda deactivate & conda deactivate & conda deactivate & conda activate base & C:\pinokio\api\wan.git\app\venv\Scripts\activate C:\pinokio\api\wan.git\app\venv && python wgp.py --multiple-images --advanced
[GGUF][llama.cpp CUDA] kernels available.
Switching to FP16 models when possible as GPU architecture doesn't support optimed BF16 Kernels
[INT8] Backend: Comfy Kitchen CUDA (setting: auto).
Downloading ffmpeg-9.0.2-essentials_build.zip: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████| 115M/115M [01:15<00:00, 1.52MB/s]
Loaded plugin: Motion Designer (from motion_designer)

██ Detached from Shell b90bdeeb-f8fd-4ae5-a925-48382f5d518b

WanGP web authentication disabled.
Deepy Web app: http://127.0.0.1:42003/deepy/

===================================================

input.event

[
"http://127.0.0.1:42003"
]

===================================================

local variables

{
"url": "http://127.0.0.1:42003"
}

  • Running on local URL: http://127.0.0.1:42003
  • To create a public link, set share=True in launch().
    wan2.1_text2video_14B_quanto_mfp16_int8.safetensors: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████| 14.9G/14.9G [1:58:30<00:00, 2.10MB/s]
    special_tokens_map.json: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 6.62k/6.62k [00:00<00:00, 2.01MB/s]
    umt5-xxl/spiece.model: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4.55M/4.55M [00:02<00:00, 1.60MB/s]
    umt5-xxl/tokenizer.json: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 16.8M/16.8M [00:03<00:00, 4.89MB/s]
    tokenizer_config.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 61.7k/61.7k [00:00<00:00, 39.9MB/s]
    xlm-roberta-large/models_clip_open-clip-xlm-roberta-large-vit-huge-14-bf16.safetensors: 100%|█████████████████████████████████████████████████████████████████| 2.39G/2.39G [15:41<00:00, 2.54MB/s]
    xlm-roberta-large/sentencepiece.bpe.model: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 5.07M/5.07M [00:04<00:00, 1.05MB/s]
    special_tokens_map.json: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 280/280 [00:00<?, ?B/s]
    xlm-roberta-large/tokenizer.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 17.1M/17.1M [00:05<00:00, 2.97MB/s]
    tokenizer_config.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 418/418 [00:00<?, ?B/s]
    Wan2.1_VAE.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 508M/508M [05:11<00:00, 1.63MB/s]
    Wan2.1_VAE_upscale2x_imageonly_real_v1.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 508M/508M [04:40<00:00, 1.81MB/s]
    Loading Model 'ckpts\wan2.1_text2video_14B_quanto_mfp16_int8.safetensors' ...
    umt5-xxl/models_t5_umt5-xxl-enc-quanto_int8.safetensors: 100%|████████████████████████████████████████████████████████████████████████████████████████████████| 6.73G/6.73G [50:16<00:00, 2.23MB/s]
    Loading Text Encoder 'ckpts\umt5-xxl\models_t5_umt5-xxl-enc-quanto_int8.safetensors' ...
    ************ Memory Management for the GPU Poor (mmgp 3.8.2) by DeepBeepMeep ************
    Switching to partial pinning since full requirements for pinned models is 13859.0 MB while estimated available reservable RAM is 13035.2 MB. You may increase the value of parameter 'perc_reserved_mem_max' to a value higher than 0.40 to force full pinnning.
    Partial pinning of data of 'transformer' to reserved RAM
    The model was partially pinned to reserved RAM: 59 large blocks spread across 12943.35 MB
    Hooked to model 'transformer' (WanModel)
    Async loading plan for model 'transformer' : base size of 2.54 MB will be preloaded with a 335.35 MB async circular shuttle
    Hooked to model 'vae' (WanVAE_)
    Hooked to model 'text_encoder' (T5Encoder)
    Async loading plan for model 'text_encoder' : base size of 2003.01 MB will be preloaded with a 184.10 MB async circular shuttle
    0%| | 0/30 [00:00<?, ?it/s][WanGP][Triton] Preparing _attn_fwd (compiling, please wait)...
    C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\attn_qk_int8_block_varlen.py:19:14: error: 'arith.extf' op operand #0 must be floating-point-like, but got 'tensor<128x128xi8, #ttg.dot_op<{opIdx = 0, parent = #ttg.blocked<{sizePerThread = [4, 4], threadsPerWarp = [2, 16], warpsPerCTA = [8, 1], order = [1, 0]}>}>>'
    qk = tl.dot(q, k).to(tl.float32) * q_scale * k_scale
    ^

C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\attn_qk_int8_block_varlen.py:95:16: note: called from
acc, l_i = _attn_fwd_inner(acc, l_i, m_i, q, q_scale, kv_len, K_ptrs, K_scale_ptr, V_ptrs, stride_kn, stride_vn,
^
module {
tt.func public @_attn_fwd(%arg0: !tt.ptr {tt.divisibility = 16 : i32}, %arg1: !tt.ptr {tt.divisibility = 16 : i32}, %arg2: !tt.ptr {tt.divisibility = 16 : i32}, %arg3: !tt.ptr {tt.divisibility = 16 : i32}, %arg4: !tt.ptr {tt.divisibility = 16 : i32}, %arg5: !tt.ptr {tt.divisibility = 16 : i32}, %arg6: !tt.ptr {tt.divisibility = 16 : i32}, %arg7: !tt.ptr {tt.divisibility = 16 : i32}, %arg8: !tt.ptr {tt.divisibility = 16 : i32}, %arg9: !tt.ptr {tt.divisibility = 16 : i32}, %arg10: i32 {tt.divisibility = 16 : i32}, %arg11: i32 {tt.divisibility = 16 : i32}, %arg12: i32 {tt.divisibility = 16 : i32}, %arg13: i32 {tt.divisibility = 16 : i32}, %arg14: i32 {tt.divisibility = 16 : i32}, %arg15: i32 {tt.divisibility = 16 : i32}, %arg16: i32 {tt.divisibility = 16 : i32}, %arg17: i32 {tt.divisibility = 16 : i32}) attributes {noinline = false} {
%cst = arith.constant dense<1.000000e+00> : tensor<128xf32>
%cst_0 = arith.constant dense<0xFF800000> : tensor<128xf32>
%c0_i32 = arith.constant 0 : i32
%c64_i32 = arith.constant 64 : i32
%cst_1 = arith.constant dense<0> : tensor<128x64xi32>
%cst_2 = arith.constant dense<0.000000e+00> : tensor<128x128xf16>
%cst_3 = arith.constant dense<0.000000e+00> : tensor<128x128xf32>
%c40_i64 = arith.constant 40 : i64
%c40_i32 = arith.constant 40 : i32
%c128_i32 = arith.constant 128 : i32
%c1_i32 = arith.constant 1 : i32
%0 = tt.get_program_id x : i32
%1 = tt.get_program_id z : i32
%2 = arith.extsi %1 : i32 to i64
%3 = tt.get_program_id y : i32
%4 = arith.extsi %3 : i32 to i64
%5 = tt.addptr %arg3, %2 : !tt.ptr, i64
%6 = tt.load %5 : !tt.ptr
%7 = tt.addptr %5, %c1_i32 : !tt.ptr, i32
%8 = tt.load %7 : !tt.ptr
%9 = arith.subi %8, %6 : i32
%10 = arith.muli %0, %c128_i32 : i32
%11 = arith.cmpi sge, %10, %9 : i32
cf.cond_br %11, ^bb1, ^bb2
^bb1: // pred: ^bb0
tt.return
^bb2: // pred: ^bb0
%12 = tt.addptr %arg7, %2 : !tt.ptr, i64
%13 = tt.load %12 : !tt.ptr
%14 = tt.addptr %arg8, %2 : !tt.ptr, i64
%15 = tt.load %14 : !tt.ptr
%16 = arith.muli %13, %c40_i64 : i64
%17 = arith.addi %16, %4 : i64
%18 = arith.muli %0, %c40_i32 : i32
%19 = arith.extsi %18 : i32 to i64
%20 = arith.addi %17, %19 : i64
%21 = arith.muli %15, %c40_i64 : i64
%22 = arith.addi %21, %4 : i64
%23 = tt.addptr %arg4, %2 : !tt.ptr, i64
%24 = tt.load %23 : !tt.ptr
%25 = tt.addptr %23, %c1_i32 : !tt.ptr, i32
%26 = tt.load %25 : !tt.ptr
%27 = arith.subi %26, %24 : i32
%28 = tt.make_range {end = 128 : i32, start = 0 : i32} : tensor<128xi32>
%29 = tt.splat %10 : i32 -> tensor<128xi32>
%30 = arith.addi %29, %28 : tensor<128xi32>
%31 = tt.make_range {end = 64 : i32, start = 0 : i32} : tensor<64xi32>
%32 = arith.muli %6, %arg11 : i32
%33 = arith.extsi %arg10 : i32 to i64
%34 = arith.muli %4, %33 : i64
%35 = arith.extsi %32 : i32 to i64
%36 = arith.addi %35, %34 : i64
%37 = tt.addptr %arg0, %36 : !tt.ptr, i64
%38 = tt.expand_dims %30 {axis = 1 : i32} : tensor<128xi32> -> tensor<128x1xi32>
%39 = tt.splat %arg11 : i32 -> tensor<128x1xi32>
%40 = arith.muli %38, %39 : tensor<128x1xi32>
%41 = tt.splat %37 : !tt.ptr -> tensor<128x1x!tt.ptr>
%42 = tt.addptr %41, %40 : tensor<128x1x!tt.ptr>, tensor<128x1xi32>
%43 = tt.expand_dims %28 {axis = 0 : i32} : tensor<128xi32> -> tensor<1x128xi32>
%44 = tt.broadcast %42 : tensor<128x1x!tt.ptr> -> tensor<128x128x!tt.ptr>
%45 = tt.broadcast %43 : tensor<1x128xi32> -> tensor<128x128xi32>
%46 = tt.addptr %44, %45 : tensor<128x128x!tt.ptr>, tensor<128x128xi32>
%47 = tt.addptr %arg5, %20 : !tt.ptr, i64
%48 = arith.muli %24, %arg13 : i32
%49 = arith.extsi %arg12 : i32 to i64
%50 = arith.muli %4, %49 : i64
%51 = arith.extsi %48 : i32 to i64
%52 = arith.addi %51, %50 : i64
%53 = tt.addptr %arg1, %52 : !tt.ptr, i64
%54 = tt.expand_dims %31 {axis = 0 : i32} : tensor<64xi32> -> tensor<1x64xi32>
%55 = tt.splat %arg13 : i32 -> tensor<1x64xi32>
%56 = arith.muli %54, %55 : tensor<1x64xi32>
%57 = tt.splat %53 : !tt.ptr -> tensor<1x64x!tt.ptr>
%58 = tt.addptr %57, %56 : tensor<1x64x!tt.ptr>, tensor<1x64xi32>
%59 = tt.expand_dims %28 {axis = 1 : i32} : tensor<128xi32> -> tensor<128x1xi32>
%60 = tt.broadcast %58 : tensor<1x64x!tt.ptr> -> tensor<128x64x!tt.ptr>
%61 = tt.broadcast %59 : tensor<128x1xi32> -> tensor<128x64xi32>
%62 = tt.addptr %60, %61 : tensor<128x64x!tt.ptr>, tensor<128x64xi32>
%63 = tt.addptr %arg6, %22 : !tt.ptr, i64
%64 = arith.muli %24, %arg15 : i32
%65 = arith.extsi %arg14 : i32 to i64
%66 = arith.muli %4, %65 : i64
%67 = arith.extsi %64 : i32 to i64
%68 = arith.addi %67, %66 : i64
%69 = tt.addptr %arg2, %68 : !tt.ptr, i64
%70 = tt.expand_dims %31 {axis = 1 : i32} : tensor<64xi32> -> tensor<64x1xi32>
%71 = tt.splat %arg15 : i32 -> tensor<64x1xi32>
%72 = arith.muli %70, %71 : tensor<64x1xi32>
%73 = tt.splat %69 : !tt.ptr -> tensor<64x1x!tt.ptr>
%74 = tt.addptr %73, %72 : tensor<64x1x!tt.ptr>, tensor<64x1xi32>
%75 = tt.broadcast %74 : tensor<64x1x!tt.ptr> -> tensor<64x128x!tt.ptr>
%76 = tt.broadcast %43 : tensor<1x128xi32> -> tensor<64x128xi32>
%77 = tt.addptr %75, %76 : tensor<64x128x!tt.ptr>, tensor<64x128xi32>
%78 = arith.muli %6, %arg17 : i32
%79 = arith.extsi %arg16 : i32 to i64
%80 = arith.muli %4, %79 : i64
%81 = arith.extsi %78 : i32 to i64
%82 = arith.addi %81, %80 : i64
%83 = tt.addptr %arg9, %82 : !tt.ptr, i64
%84 = tt.splat %arg17 : i32 -> tensor<128x1xi32>
%85 = arith.muli %38, %84 : tensor<128x1xi32>
%86 = tt.splat %83 : !tt.ptr -> tensor<128x1x!tt.ptr>
%87 = tt.addptr %86, %85 : tensor<128x1x!tt.ptr>, tensor<128x1xi32>
%88 = tt.broadcast %87 : tensor<128x1x!tt.ptr> -> tensor<128x128x!tt.ptr>
%89 = tt.addptr %88, %45 : tensor<128x128x!tt.ptr>, tensor<128x128xi32>
%90 = tt.splat %9 : i32 -> tensor<128x1xi32>
%91 = arith.cmpi slt, %38, %90 : tensor<128x1xi32>
%92 = tt.broadcast %91 : tensor<128x1xi1> -> tensor<128x128xi1>
%93 = tt.load %46, %92 : tensor<128x128x!tt.ptr>
%94 = tt.load %47 : !tt.ptr
%95:6 = scf.for %arg18 = %c0_i32 to %27 step %c64_i32 iter_args(%arg19 = %cst_3, %arg20 = %cst, %arg21 = %cst_0, %arg22 = %62, %arg23 = %63, %arg24 = %77) -> (tensor<128x128xf32>, tensor<128xf32>, tensor<128xf32>, tensor<128x64x!tt.ptr>, !tt.ptr, tensor<64x128x!tt.ptr>) : i32 {
%100 = arith.subi %27, %arg18 : i32
%101 = tt.splat %100 : i32 -> tensor<1x64xi32>
%102 = arith.cmpi slt, %54, %101 : tensor<1x64xi32>
%103 = tt.broadcast %102 : tensor<1x64xi1> -> tensor<128x64xi1>
%104 = tt.load %arg22, %103 : tensor<128x64x!tt.ptr>
%105 = tt.load %arg23 : !tt.ptr
%106 = tt.dot %93, %104, %cst_1 : tensor<128x128xi8> * tensor<128x64xi8> -> tensor<128x64xi32>
%107 = arith.sitofp %106 : tensor<128x64xi32> to tensor<128x64xf32>
%108 = tt.splat %94 : f32 -> tensor<128x64xf32>
%109 = arith.mulf %107, %108 : tensor<128x64xf32>
%110 = tt.splat %105 : f32 -> tensor<128x64xf32>
%111 = arith.mulf %109, %110 : tensor<128x64xf32>
%112 = "tt.reduce"(%111) <{axis = 1 : i32}> ({
^bb0(%arg25: f32, %arg26: f32):
%141 = arith.maxnumf %arg25, %arg26 : f32
tt.reduce.return %141 : f32
}) : (tensor<128x64xf32>) -> tensor<128xf32>
%113 = arith.maxnumf %arg21, %112 : tensor<128xf32>
%114 = tt.expand_dims %113 {axis = 1 : i32} : tensor<128xf32> -> tensor<128x1xf32>
%115 = tt.broadcast %114 : tensor<128x1xf32> -> tensor<128x64xf32>
%116 = arith.subf %111, %115 : tensor<128x64xf32>
%117 = math.exp2 %116 : tensor<128x64xf32>
%118 = "tt.reduce"(%117) <{axis = 1 : i32}> ({
^bb0(%arg25: f32, %arg26: f32):
%141 = arith.addf %arg25, %arg26 : f32
tt.reduce.return %141 : f32
}) : (tensor<128x64xf32>) -> tensor<128xf32>
%119 = arith.subf %arg21, %113 : tensor<128xf32>
%120 = math.exp2 %119 : tensor<128xf32>
%121 = arith.mulf %arg20, %120 : tensor<128xf32>
%122 = arith.addf %121, %118 : tensor<128xf32>
%123 = tt.expand_dims %120 {axis = 1 : i32} : tensor<128xf32> -> tensor<128x1xf32>
%124 = tt.broadcast %123 : tensor<128x1xf32> -> tensor<128x128xf32>
%125 = arith.mulf %arg19, %124 : tensor<128x128xf32>
%126 = tt.splat %100 : i32 -> tensor<64x1xi32>
%127 = arith.cmpi slt, %70, %126 : tensor<64x1xi32>
%128 = tt.broadcast %127 : tensor<64x1xi1> -> tensor<64x128xi1>
%129 = tt.load %arg24, %128 : tensor<64x128x!tt.ptr>
%130 = arith.truncf %117 : tensor<128x64xf32> to tensor<128x64xf16>
%131 = tt.dot %130, %129, %cst_2 : tensor<128x64xf16> * tensor<64x128xf16> -> tensor<128x128xf16>
%132 = arith.extf %131 : tensor<128x128xf16> to tensor<128x128xf32>
%133 = arith.addf %125, %132 : tensor<128x128xf32>
%134 = arith.muli %arg13, %c64_i32 : i32
%135 = tt.splat %134 : i32 -> tensor<128x64xi32>
%136 = tt.addptr %arg22, %135 : tensor<128x64x!tt.ptr>, tensor<128x64xi32>
%137 = tt.addptr %arg23, %c40_i32 : !tt.ptr, i32
%138 = arith.muli %arg15, %c64_i32 : i32
%139 = tt.splat %138 : i32 -> tensor<64x128xi32>
%140 = tt.addptr %arg24, %139 : tensor<64x128x!tt.ptr>, tensor<64x128xi32>
scf.yield %133, %122, %113, %136, %137, %140 : tensor<128x128xf32>, tensor<128xf32>, tensor<128xf32>, tensor<128x64x!tt.ptr>, !tt.ptr, tensor<64x128x!tt.ptr>
}
%96 = tt.expand_dims %95#1 {axis = 1 : i32} : tensor<128xf32> -> tensor<128x1xf32>
%97 = tt.broadcast %96 : tensor<128x1xf32> -> tensor<128x128xf32>
%98 = arith.divf %95#0, %97 : tensor<128x128xf32>
%99 = arith.truncf %98 : tensor<128x128xf32> to tensor<128x128xf16>
tt.store %89, %99, %92 : tensor<128x128x!tt.ptr>
tt.return
}
}

{-#
external_resources: {
mlir_reproducer: {
pipeline: "builtin.module(convert-triton-to-tritongpu{enable-source-remat=false num-ctas=1 num-warps=8 target=cuda:75 threads-per-warp=32}, tritongpu-coalesce, tritongpu-F32DotTC{emu-tf32=false}, triton-nvidia-gpu-plan-cta, tritongpu-remove-layout-conversions, tritongpu-optimize-thread-locality, tritongpu-accelerate-matmul, tritongpu-remove-layout-conversions, tritongpu-optimize-dot-operands{hoist-layout-conversion=false}, triton-nvidia-optimize-descriptor-encoding, triton-loop-aware-cse, triton-licm, canonicalize{cse-between-iterations=false max-iterations=10 max-num-rewrites=-1 region-simplify=normal test-convergence=false top-down=true}, triton-loop-aware-cse, tritongpu-optimize-dot-operands{hoist-layout-conversion=false}, tritongpu-coalesce-async-copy, triton-nvidia-optimize-tmem-layouts, triton-nvidia-tmem-load-reduce, tritongpu-remove-layout-conversions, triton-nvidia-interleave-tmem, tritongpu-reduce-data-duplication, tritongpu-reorder-instructions, triton-loop-aware-cse, symbol-dce, triton-nvidia-gpu-fence-insertion{compute-capability=75}, triton-nvidia-mma-lowering, sccp, cse, canonicalize{cse-between-iterations=false max-iterations=10 max-num-rewrites=-1 region-simplify=normal test-convergence=false top-down=true})",
disable_threading: true,
verify_each: true
}
}
#-}
C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\attn_qk_int8_block_varlen.py:41:1: error: Failures have been detected while processing an MLIR pass pipeline
def _attn_fwd(Q, K, V,
^
C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\attn_qk_int8_block_varlen.py:41:1: note: Pipeline failed while executing [TritonGPUAccelerateMatmul on 'builtin.module' operation]: reproducer generated at std::errs, please share the reproducer above with Triton project.
0%| | 0/30 [00:02<?, ?it/s]
Traceback (most recent call last):
File "C:\pinokio\api\wan.git\app\wgp.py", line 8088, in generate_media
samples = wan_model.generate(
^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\shared\utils\phase_progress.py", line 53, in wrapped
result = method(self, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\models\wan\any2video.py", line 1751, in generate
noise_pred = denoise_fn(latents)
^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\models\wan\any2video.py", line 1658, in denoise_with_cfg_fn
ret_values = trans( **gen_args , **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1776, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1787, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\mmgp\offload.py", line 3399, in check_change_module
return previous_method(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\models\wan\modules\model.py", line 2036, in forward
x_list[i] = block(x, context = context, hints= hints, audio_scale= audio_scale, multitalk_audio = multitalk_audio, multitalk_masks =multitalk_masks, e= e0, motion_vec = motion_vec, lynx_ip_embeds= lynx_ip_embeds, lynx_ref_buffer = lynx_ref_buffer, sub_x_no =i, **block_kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1776, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1787, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\mmgp\offload.py", line 3377, in check_load_into_GPU_needed_other
return previous_method(*args, **kwargs) # other
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\models\wan\modules\model.py", line 653, in forward
y, x_ref_attn_map = self.self_attn( xlist, grid_sizes, freqs, block_mask = block_mask, ref_target_masks = multitalk_masks, ref_images_count = ref_images_count, standin_phase= standin_phase, lynx_ref_buffer = lynx_ref_buffer, lynx_ref_scale = lynx_ref_scale, sub_x_no = sub_x_no)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1776, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch\nn\modules\module.py", line 1787, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\mmgp\offload.py", line 3377, in check_load_into_GPU_needed_other
return previous_method(*args, **kwargs) # other
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\models\wan\modules\model.py", line 385, in forward
x = pay_attention( qkv_list, recycle_q=True)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\torch_dynamo\eval_frame.py", line 1181, in _fn
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\shared\attention.py", line 509, in pay_attention
x = sageattn_varlen_wrapper(
^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\shared\attention.py", line 68, in sageattn_varlen_wrapper
return sageattn_varlen(
^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\core.py", line 203, in sageattn_varlen
o = attn_false_varlen(q_int8, k_int8, v, cu_seqlens_q, cu_seqlens_k, max_seqlen_q, q_scale, k_scale, cu_seqlens_q_scale, cu_seqlens_k_scale, output_dtype=dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\sageattention\attn_qk_int8_block_varlen.py", line 119, in forward
_attn_fwd[grid](
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\runtime\jit.py", line 374, in
return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\runtime\jit.py", line 757, in run
kernel = self._do_compile(key, signature, device, constexprs, options, attrs, warmup)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\runtime\jit.py", line 902, in _do_compile
kernel = self.compile(src, target=target, options=options.dict)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\compiler\compiler.py", line 327, in compile
next_module = compile_ir(module, metadata)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\backends\nvidia\compiler.py", line 611, in
stages["ttgir"] = lambda src, metadata: self.make_ttgir(src, metadata, options, capability)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\pinokio\api\wan.git\app\venv\Lib\site-packages\triton\backends\nvidia\compiler.py", line 346, in make_ttgir
pm.run(mod, 'make_ttgir')
RuntimeError: PassManager::run failed
Error Queue autosaved successfully to error_queue.zip

AboutWan2GPpinokiofactory/wanView app
Loading conversation…

Replies

Post context

App used

Wan2GPpinokiofactory/wan

1-click WanGP Launcher. Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Vide...

View app

Published by

@designer_992h ago
Pinokio
PrivacyTerms
FeedApps

Post context

App used

Wan2GPpinokiofactory/wan

1-click WanGP Launcher. Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Vide...

View app

Published by

@designer_992h ago