Describe the bug
A Z-Image LoRA that only carries the model's original attention module names, a fused attention_qkv and a bare attention_out per block (what musubi-tuner and LyCORIS emit when trained against the original Z-Image module layout), cannot be loaded: _convert_non_diffusers_z_image_lora_to_diffusers drops both keys on the assumption that split to.q/k/v and to_out.0 keys are also present (the Anime-Z layout, where they are redundant), which empties the LoRA, and then raises ValueError: state_dict should be empty on the leftover .alpha keys of the out modules.
Expected: the fused qkv LoRA is split into to_q / to_k / to_v (chunks of the up weight along dim 0, like convert_z_image_fused_attention does for base weights) and out maps to to_out.0.
Reproduction
import torch
from diffusers.loaders.lora_conversion_utils import _convert_non_diffusers_z_image_lora_to_diffusers
rank, dim = 2, 8
sd = {}
for block in ("layers_0", "context_refiner_0"):
sd[f"lora_unet_{block}_attention_qkv.lora_down.weight"] = torch.randn(rank, dim)
sd[f"lora_unet_{block}_attention_qkv.lora_up.weight"] = torch.randn(3 * dim, rank)
sd[f"lora_unet_{block}_attention_qkv.alpha"] = torch.tensor(float(rank))
sd[f"lora_unet_{block}_attention_out.lora_down.weight"] = torch.randn(rank, dim)
sd[f"lora_unet_{block}_attention_out.lora_up.weight"] = torch.randn(dim, rank)
sd[f"lora_unet_{block}_attention_out.alpha"] = torch.tensor(float(rank))
_convert_non_diffusers_z_image_lora_to_diffusers(sd)
The same happens end to end through ZImagePipeline.load_lora_weights with a real file in that layout (dims 3840 → 11520 for qkv, 3840 → 3840 for out, 34 blocks).
Logs
ValueError: `state_dict` should be empty at this point but has state_dict.keys()=dict_keys(['context_refiner.0.attention.to_out.0.alpha', 'layers.0.attention.to_out.0.alpha'])
Without the .alpha keys the call succeeds and returns an empty dict: the whole LoRA is silently dropped.
System Info
diffusers main at e0118ad (also 0.40.0), torch 2.14.0, Python 3.12, Linux.
Who can help?
@sayakpaul @BenjaminBossan
Fix in #14875.
Describe the bug
A Z-Image LoRA that only carries the model's original attention module names, a fused
attention_qkvand a bareattention_outper block (what musubi-tuner and LyCORIS emit when trained against the original Z-Image module layout), cannot be loaded:_convert_non_diffusers_z_image_lora_to_diffusersdrops both keys on the assumption that splitto.q/k/vandto_out.0keys are also present (the Anime-Z layout, where they are redundant), which empties the LoRA, and then raisesValueError: state_dict should be emptyon the leftover.alphakeys of theoutmodules.Expected: the fused
qkvLoRA is split intoto_q/to_k/to_v(chunks of the up weight along dim 0, likeconvert_z_image_fused_attentiondoes for base weights) andoutmaps toto_out.0.Reproduction
The same happens end to end through
ZImagePipeline.load_lora_weightswith a real file in that layout (dims 3840 → 11520 forqkv, 3840 → 3840 forout, 34 blocks).Logs
Without the
.alphakeys the call succeeds and returns an empty dict: the whole LoRA is silently dropped.System Info
diffusers main at e0118ad (also 0.40.0), torch 2.14.0, Python 3.12, Linux.
Who can help?
@sayakpaul @BenjaminBossan
Fix in #14875.