FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Per-module alphas of kohya text encoder LoRAs collapse to the most common alpha (alpha_pattern keys stop at the parent module) · Issue #14970 · huggingface/diffusers · GitHub

Repository navigation

Per-module alphas of kohya text encoder LoRAs collapse to the most common alpha (alpha_pattern keys stop at the parent module) #14970

Description

Describe the bug

When a kohya-format SDXL/SD1.5 LoRA has different .alpha values across text encoder modules, every text encoder module ends up with the most common alpha. The minority modules get the wrong scale.

get_peft_kwargs builds the text encoder alpha_pattern with ".".join(k.split(".down.")[0].split(".")[:-1]) (peft_utils.py#L189). For a network alpha key such as text_model.encoder.layers.1.self_attn.to_k_lora.down.weight.alpha, that yields text_model.encoder.layers.1.self_attn, which is the parent attention block and not the k_proj module. PEFT's get_pattern_key matches pattern keys against the end of the module name, so the key matches no LoRA module and the default lora_alpha (the most common value) applies everywhere.

The UNet branch keeps the module name, so the UNet is not affected. The same result occurs with a text encoder that still has the text_model level (e.g. CLIPTextModelWithProjection), so this is independent of #14872.

For scale: in a library of ~8.9k SDXL kohya LoRAs, 59 have per-module text encoder alphas (values from 0.5 to 6 at the same rank, mostly character LoRAs), so their text encoder modules can be off by up to 12x.

Reproduction

import torch
from diffusers import StableDiffusionXLPipeline
from transformers import CLIPTextConfig, CLIPTextModel

config = CLIPTextConfig(
    hidden_size=8, intermediate_size=16, num_hidden_layers=2, num_attention_heads=2,
    vocab_size=50, max_position_embeddings=8, bos_token_id=0, eos_token_id=1, pad_token_id=1,
)
text_encoder = CLIPTextModel(config)

rank = 4

def kohya_module(name, alpha):
    return {
        f"lora_te1_text_model_{name}.lora_down.weight": torch.ones(rank, 8),
        f"lora_te1_text_model_{name}.lora_up.weight": torch.ones(8, rank),
        f"lora_te1_text_model_{name}.alpha": torch.tensor(alpha),
    }

# Two modules with alpha 1.0 (the most common value) and one with alpha 4.0.
lora = {
    **kohya_module("encoder_layers_0_self_attn_q_proj", 1.0),
    **kohya_module("encoder_layers_0_self_attn_out_proj", 1.0),
    **kohya_module("encoder_layers_1_self_attn_k_proj", 4.0),
}

state_dict, network_alphas = StableDiffusionXLPipeline.lora_state_dict(lora)
StableDiffusionXLPipeline.load_lora_into_text_encoder(
    state_dict, network_alphas, text_encoder=text_encoder, prefix="text_encoder", adapter_name="test"
)

k_proj = text_encoder.encoder.layers[1].self_attn.k_proj
print("k_proj scaling:", k_proj.scaling["test"], "expected:", 4.0 / rank)
print("peft alpha_pattern:", text_encoder.peft_config["test"].alpha_pattern)

Logs

k_proj scaling: 0.25 expected: 1.0
peft alpha_pattern: {'text_model.encoder.layers.1.self_attn': 4.0}

System Info

  • 🤗 Diffusers version: 0.41.0
  • Platform: Linux (WSL2) x86_64
  • Python version: 3.14.3
  • PyTorch version (GPU?): 2.13.0+cu130 (True)
  • Huggingface_hub version: 1.33.0
  • Transformers version: 5.16.1
  • Accelerate version: 1.15.0
  • PEFT version: 0.20.0
  • Safetensors version: 0.8.0
  • Using GPU in script?: No
  • Using distributed or parallel set-up in script?: No

Who can help?

@sayakpaul @BenjaminBossan

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL