Model Config Attribute Error
2025-02-022 turns8,031 charsgpt-4o-mini
Summary
User encountered a Python error while running main.py related to model configuration attributes.
Messages
python3 main.py
Python version is above 3.10, patching the collections module.
/Users/danielkliewer/Janus-main/venv/lib/python3.10/site-packages/tra
nsformers/models/auto/image_processing_auto.py:590: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead warnings.warn(
Using a slow image processor as `use_fast` is unset and a slow proces
sor was saved with this model. `use_fast=True` will be the default behavior in v4.48, even if the model was saved with a slow processor. This will result in minor differences in outputs. You'll still be able to use a slow processor with `use_fast=False`. You are using the default legacy behaviour of the <class 'transformer
s.models.llama.tokenization_llama_fast.LlamaTokenizerFast'>. This is expected, and simply means that the `legacy` (previous) behavior will be used so nothing changes for you. If you want to use the new behaviour, set `legacy=False`. This should only be set if you understand what it means, and thoroughly read the reason why this was added as explained in https://github.com/huggingface/transformers/pull/24565 - if you loaded a llama tokenizer from a GGUF file you can ignore this message. Some kwargs in processor config are unused and will not have any effe
ct: num_image_tokens, add_special_token, sft_format, image_tag, ignore_id, mask_prompt. Loading checkpoint shards: 100%|███████| 2/2 [00:08<00:00, 4.11s/it]
Traceback (most recent call last):
File "/Users/danielkliewer/Janus-main/main.py", line 25, in <module
> num_heads = vl_gpt.config.num_attention_heads # This is the crit
ical fix File "/Users/danielkliewer/Janus-main/venv/lib/python3.10/site-pack
ages/transformers/configuration_utils.py", line 211, in __getattribute__ return super().__getattribute__(key)
AttributeError: 'MultiModalityConfig' object has no attribute 'num_at
tention_heads' import os
import torch.nn as nn
import PIL.Image
import torch
import numpy as np
from transformers import AutoModelForCausalLM
from janus.models import MultiModalityCausalLM, VLChatProcessor
# Load Model and Processor
model_path = "Janus-Pro-7B"
vl_chat_processor: VLChatProcessor = VLChatProcessor.from_pretrained(model_path)
tokenizer = vl_chat_processor.tokenizer
device = torch.device("mps") # Apple Silicon GPU
dtype = torch.float16 # Ensure everything runs in float16
# Projection Layer - Move this after device definition
projection_layer = nn.Linear(630784, 4096).to(device, dtype=torch.float16)
vl_gpt: MultiModalityCausalLM = AutoModelForCausalLM.from_pretrained(
model_path, trust_remote_code=True
).to(dtype).to(device).eval()
# Get actual number of attention heads from model config
num_heads = vl_gpt.config.num_attention_heads # This is the critical fix
# Image Generation Parameters
TEMPERATURE = 0.7
CFG_WEIGHT = 7
IMAGE_TOKEN_NUM = 1024
IMG_SIZE = 512
PATCH_SIZE = 16
conversation = [
{
"role": "<|User|>",
"content": "A black and white domestic shorthair cat sits at a futuristic cyberpunk-style control station. The cat’s neon-blue eyes emit a faint glow. It is positioned in front of a high-tech mechanical keyboard with illuminated keys, its paws resting on the surface. Surrounding the cat are multiple floating holographic screens displaying scrolling lines of code, world maps, and red security breach alerts. The environment is dimly lit with neon lighting—deep crimson and electric blue reflections cast onto metallic surfaces. Large server racks line the background, emitting a soft electronic glow. A massive ultra-wide monitor behind the cat displays cryptic data streams. The atmosphere is cinematic, high-tech, and futuristic, emphasizing the cyberpunk aesthetic.",
},
{"role": "<|Assistant|>", "content": ""},
]
sft_format = vl_chat_processor.apply_sft_template_for_multi_turn_prompts(
conversations=conversation,
sft_format=vl_chat_processor.sft_format,
system_prompt="",
)
prompt = sft_format + vl_chat_processor.image_start_tag
@torch.inference_mode()
def generate(mmgpt: MultiModalityCausalLM, processor: VLChatProcessor, prompt: str):
input_ids = processor.tokenizer(prompt, return_tensors="pt").input_ids.to(device, dtype=torch.long)
inputs_embeds = mmgpt.language_model.get_input_embeddings()(input_ids).to(dtype)
generated_tokens = torch.zeros((1, IMAGE_TOKEN_NUM), dtype=torch.int).to(device)
past_key_values = None
for i in range(IMAGE_TOKEN_NUM):
outputs = mmgpt.language_model.model(
inputs_embeds=inputs_embeds,
use_cache=True,
past_key_values=past_key_values
)
hidden_states = outputs.last_hidden_state.to(dtype)
# Use model's actual number of attention heads
head_dim = hidden_states.shape[-1] // num_heads
hidden_states = hidden_states.view(1, num_heads, -1, head_dim) # Correct reshaping
hidden_states = hidden_states.view(1, -1)
hidden_states = projection_layer(hidden_states)
past_key_values = outputs.past_key_values
logits = mmgpt.gen_head(hidden_states).to(dtype)
logits = logits + CFG_WEIGHT * (logits - logits)
probs = torch.softmax(logits / TEMPERATURE, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
generated_tokens[:, i] = next_token.squeeze(dim=-1)
img_embeds = mmgpt.prepare_gen_img_embeds(next_token).to(dtype)
inputs_embeds = img_embeds.unsqueeze(dim=1)
decoded_image = mmgpt.gen_vision_model.decode_code(
generated_tokens.to(dtype=torch.int),
shape=[1, 8, IMG_SIZE // PATCH_SIZE, IMG_SIZE // PATCH_SIZE]
).to(torch.float32).cpu().numpy().transpose(0, 2, 3, 1)
decoded_image = ((decoded_image - decoded_image.min()) / (decoded_image.max() - decoded_image.min()) * 255).astype(np.uint8)
os.makedirs('generated_samples', exist_ok=True)
save_path = os.path.join('generated_samples', "img.jpg")
PIL.Image.fromarray(decoded_image[0]).save(save_path)
print(f"Image saved to {save_path}")
generate(vl_gpt, vl_chat_processor, prompt)
The error you're encountering is related to the model's configuration. Specifically, `MultiModalityConfig` doesn't have the attribute `num_attention_heads`, which is causing the `AttributeError`.
To address this, you can check the model's configuration for available attributes. It's likely that the `num_attention_heads` attribute might be part of a different configuration class or nested inside another attribute.
To resolve the issue:
1. **Inspect the Model Configuration**: Print out the model's configuration to check what attributes are available:
```python
print(vl_gpt.config)
```
2. **Check for Alternative Attributes**: If `num_attention_heads` is not available, search the configuration for any other related attributes like `num_layers`, `hidden_size`, or similar.
3. **Handle Missing Attributes Gracefully**: If `num_attention_heads` is absent, consider using default values or adapting the logic to work with available configuration parameters.
4. **Update Code for Compatibility**: If the model architecture has changed (e.g., multi-modality models might not directly expose the same attributes as GPT-like models), ensure that the handling of tokenization, embeddings, and attention mechanisms is appropriate for this type of model.
Once you've identified the correct attribute or structure, you can update your code accordingly.
Let me know if you'd like further assistance with debugging!
gpt-4o-mini