← all conversations

Model Config Attribute Error

2025-02-022 turns8,031 charsgpt-4o-mini
pythonmodel-configurationattribute-error

Summary

User encountered a Python error while running main.py related to model configuration attributes.

Messages

python3 main.py Python version is above 3.10, patching the collections module. /Users/danielkliewer/Janus-main/venv/lib/python3.10/site-packages/tra nsformers/models/auto/image_processing_auto.py:590: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead warnings.warn( Using a slow image processor as `use_fast` is unset and a slow proces sor was saved with this model. `use_fast=True` will be the default behavior in v4.48, even if the model was saved with a slow processor. This will result in minor differences in outputs. You'll still be able to use a slow processor with `use_fast=False`. You are using the default legacy behaviour of the <class 'transformer s.models.llama.tokenization_llama_fast.LlamaTokenizerFast'>. This is expected, and simply means that the `legacy` (previous) behavior will be used so nothing changes for you. If you want to use the new behaviour, set `legacy=False`. This should only be set if you understand what it means, and thoroughly read the reason why this was added as explained in https://github.com/huggingface/transformers/pull/24565 - if you loaded a llama tokenizer from a GGUF file you can ignore this message. Some kwargs in processor config are unused and will not have any effe ct: num_image_tokens, add_special_token, sft_format, image_tag, ignore_id, mask_prompt. Loading checkpoint shards: 100%|███████| 2/2 [00:08<00:00, 4.11s/it] Traceback (most recent call last): File "/Users/danielkliewer/Janus-main/main.py", line 25, in <module > num_heads = vl_gpt.config.num_attention_heads # This is the crit ical fix File "/Users/danielkliewer/Janus-main/venv/lib/python3.10/site-pack ages/transformers/configuration_utils.py", line 211, in __getattribute__ return super().__getattribute__(key) AttributeError: 'MultiModalityConfig' object has no attribute 'num_at tention_heads' import os import torch.nn as nn import PIL.Image import torch import numpy as np from transformers import AutoModelForCausalLM from janus.models import MultiModalityCausalLM, VLChatProcessor # Load Model and Processor model_path = "Janus-Pro-7B" vl_chat_processor: VLChatProcessor = VLChatProcessor.from_pretrained(model_path) tokenizer = vl_chat_processor.tokenizer device = torch.device("mps") # Apple Silicon GPU dtype = torch.float16 # Ensure everything runs in float16 # Projection Layer - Move this after device definition projection_layer = nn.Linear(630784, 4096).to(device, dtype=torch.float16) vl_gpt: MultiModalityCausalLM = AutoModelForCausalLM.from_pretrained( model_path, trust_remote_code=True ).to(dtype).to(device).eval() # Get actual number of attention heads from model config num_heads = vl_gpt.config.num_attention_heads # This is the critical fix # Image Generation Parameters TEMPERATURE = 0.7 CFG_WEIGHT = 7 IMAGE_TOKEN_NUM = 1024 IMG_SIZE = 512 PATCH_SIZE = 16 conversation = [ { "role": "<|User|>", "content": "A black and white domestic shorthair cat sits at a futuristic cyberpunk-style control station. The cat’s neon-blue eyes emit a faint glow. It is positioned in front of a high-tech mechanical keyboard with illuminated keys, its paws resting on the surface. Surrounding the cat are multiple floating holographic screens displaying scrolling lines of code, world maps, and red security breach alerts. The environment is dimly lit with neon lighting—deep crimson and electric blue reflections cast onto metallic surfaces. Large server racks line the background, emitting a soft electronic glow. A massive ultra-wide monitor behind the cat displays cryptic data streams. The atmosphere is cinematic, high-tech, and futuristic, emphasizing the cyberpunk aesthetic.", }, {"role": "<|Assistant|>", "content": ""}, ] sft_format = vl_chat_processor.apply_sft_template_for_multi_turn_prompts( conversations=conversation, sft_format=vl_chat_processor.sft_format, system_prompt="", ) prompt = sft_format + vl_chat_processor.image_start_tag @torch.inference_mode() def generate(mmgpt: MultiModalityCausalLM, processor: VLChatProcessor, prompt: str): input_ids = processor.tokenizer(prompt, return_tensors="pt").input_ids.to(device, dtype=torch.long) inputs_embeds = mmgpt.language_model.get_input_embeddings()(input_ids).to(dtype) generated_tokens = torch.zeros((1, IMAGE_TOKEN_NUM), dtype=torch.int).to(device) past_key_values = None for i in range(IMAGE_TOKEN_NUM): outputs = mmgpt.language_model.model( inputs_embeds=inputs_embeds, use_cache=True, past_key_values=past_key_values ) hidden_states = outputs.last_hidden_state.to(dtype) # Use model's actual number of attention heads head_dim = hidden_states.shape[-1] // num_heads hidden_states = hidden_states.view(1, num_heads, -1, head_dim) # Correct reshaping hidden_states = hidden_states.view(1, -1) hidden_states = projection_layer(hidden_states) past_key_values = outputs.past_key_values logits = mmgpt.gen_head(hidden_states).to(dtype) logits = logits + CFG_WEIGHT * (logits - logits) probs = torch.softmax(logits / TEMPERATURE, dim=-1) next_token = torch.multinomial(probs, num_samples=1) generated_tokens[:, i] = next_token.squeeze(dim=-1) img_embeds = mmgpt.prepare_gen_img_embeds(next_token).to(dtype) inputs_embeds = img_embeds.unsqueeze(dim=1) decoded_image = mmgpt.gen_vision_model.decode_code( generated_tokens.to(dtype=torch.int), shape=[1, 8, IMG_SIZE // PATCH_SIZE, IMG_SIZE // PATCH_SIZE] ).to(torch.float32).cpu().numpy().transpose(0, 2, 3, 1) decoded_image = ((decoded_image - decoded_image.min()) / (decoded_image.max() - decoded_image.min()) * 255).astype(np.uint8) os.makedirs('generated_samples', exist_ok=True) save_path = os.path.join('generated_samples', "img.jpg") PIL.Image.fromarray(decoded_image[0]).save(save_path) print(f"Image saved to {save_path}") generate(vl_gpt, vl_chat_processor, prompt)
The error you're encountering is related to the model's configuration. Specifically, `MultiModalityConfig` doesn't have the attribute `num_attention_heads`, which is causing the `AttributeError`. To address this, you can check the model's configuration for available attributes. It's likely that the `num_attention_heads` attribute might be part of a different configuration class or nested inside another attribute. To resolve the issue: 1. **Inspect the Model Configuration**: Print out the model's configuration to check what attributes are available: ```python print(vl_gpt.config) ``` 2. **Check for Alternative Attributes**: If `num_attention_heads` is not available, search the configuration for any other related attributes like `num_layers`, `hidden_size`, or similar. 3. **Handle Missing Attributes Gracefully**: If `num_attention_heads` is absent, consider using default values or adapting the logic to work with available configuration parameters. 4. **Update Code for Compatibility**: If the model architecture has changed (e.g., multi-modality models might not directly expose the same attributes as GPT-like models), ensure that the handling of tokenization, embeddings, and attention mechanisms is appropriate for this type of model. Once you've identified the correct attribute or structure, you can update your code accordingly. Let me know if you'd like further assistance with debugging!
gpt-4o-mini