้กน็›ฎๆ–‡ไปถๅคน

ๆ–‡ไปถ
wehub-resource-sync e06fe8e8c6
Secret Leaks / trufflehog (push) Failing after 1s
Build documentation / build (push) Failing after 1s
Build documentation / build_other_lang (push) Failing after 0s
CodeQL Security Analysis / CodeQL Analysis (push) Failing after 0s
PR CI / pr-ci (push) Failing after 1s
Slow tests on important models (on Push - A10) / Get all modified files (push) Failing after 1s
Slow tests on important models (on Push - A10) / Model CI (push) Has been skipped
Self-hosted runner (benchmark) / Benchmark (aws-g5-4xlarge-cache) (push) Has been cancelled
New model PR merged notification / Notify new model (push) Has been cancelled
Update Transformers metadata / build_and_package (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 11:57:37 +08:00

12 KiB

Qwen2-VLQwen2-VL

PyTorch FlashAttention

OverviewOverview

Qwen2-VL ๋ชจ๋ธ์€ ์•Œ๋ฆฌ๋ฐ”๋ฐ” ๋ฆฌ์„œ์น˜์˜ QwenํŒ€์—์„œ ๊ฐœ๋ฐœํ•œ Qwen-VL ๋ชจ๋ธ์˜ ์ฃผ์š” ์—…๋ฐ์ดํŠธ ๋ฒ„์ „์ž…๋‹ˆ๋‹ค.

๋ธ”๋กœ๊ทธ์˜ ์š”์•ฝ์€ ๋‹ค์Œ๊ณผ ๊ฐ™์Šต๋‹ˆ๋‹ค:

์ด ๋ธ”๋กœ๊ทธ๋Š” ์ง€๋‚œ ๋ช‡ ๋…„๊ฐ„ Qwen-VL์—์„œ ์ค‘๋Œ€ํ•œ ๊ฐœ์„ ์„ ๊ฑฐ์ณ ๋ฐœ์ „๋œ Qwen2-VL ๋ชจ๋ธ์„ ์†Œ๊ฐœํ•ฉ๋‹ˆ๋‹ค. ์ค‘์š” ๊ฐœ์„  ์‚ฌํ•ญ์€ ํ–ฅ์ƒ๋œ ์ด๋ฏธ์ง€ ์ดํ•ด, ๊ณ ๊ธ‰ ๋น„๋””์˜ค ์ดํ•ด, ํ†ตํ•ฉ ์‹œ๊ฐ ์—์ด์ „ํŠธ ๊ธฐ๋Šฅ, ํ™•์žฅ๋œ ๋‹ค์–ธ์–ด ์ง€์›์„ ํฌํ•จํ•˜๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค.๋ชจ๋ธ ์•„ํ‚คํ…์ฒ˜๋Š” Naive Dynamic Resolution ์ง€์›์„ ํ†ตํ•ด ์ž„์˜์˜ ์ด๋ฏธ์ง€ ํ•ด์ƒ๋„๋ฅผ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ๋„๋ก ์ตœ์ ํ™”๋˜์—ˆ์œผ๋ฉฐ, ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ํšŒ์ „ ์œ„์น˜ ์ž„๋ฒ ๋”ฉ(M-ROPE)์„ ํ™œ์šฉํ•˜์—ฌ 1D ํ…์ŠคํŠธ์™€ ๋‹ค์ฐจ์› ์‹œ๊ฐ ๋ฐ์ดํ„ฐ๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ์ฒ˜๋ฆฌํ•ฉ๋‹ˆ๋‹ค. ์ด ์—…๋ฐ์ดํŠธ๋œ ๋ชจ๋ธ์€ ์‹œ๊ฐ ๊ด€๋ จ ์ž‘์—…์—์„œ GPT-4o์™€ Claude 3.5 Sonnet ๊ฐ™์€ ์„ ๋„์ ์ธ AI ์‹œ์Šคํ…œ๊ณผ ๊ฒฝ์Ÿ๋ ฅ ์žˆ๋Š” ์„ฑ๋Šฅ์„ ๋ณด์—ฌ์ฃผ๋ฉฐ, ํ…์ŠคํŠธ ๋Šฅ๋ ฅ์—์„œ๋Š” ์˜คํ”ˆ์†Œ์Šค ๋ชจ๋ธ ์ค‘ ์ƒ์œ„๊ถŒ์— ๋žญํฌ๋˜์–ด ์žˆ์Šต๋‹ˆ๋‹ค. ์ด๋Ÿฌํ•œ ๋ฐœ์ „์€ Qwen2-VL์„ ๊ฐ•๋ ฅํ•œ ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ์ฒ˜๋ฆฌ ๋ฐ ์ถ”๋ก  ๋Šฅ๋ ฅ์ด ํ•„์š”ํ•œ ๋‹ค์–‘ํ•œ ์‘์šฉ ๋ถ„์•ผ์—์„œ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ๋Š” ๋‹ค์žฌ๋‹ค๋Šฅํ•œ ๋„๊ตฌ๋กœ ๋งŒ๋“ค์–ด์ค๋‹ˆ๋‹ค.

drawing

Qwen2-VL ๊ตฌ์กฐ. ์ถœ์ฒ˜: ๋ธ”๋กœ๊ทธ ๊ฒŒ์‹œ๊ธ€

์ด ๋ชจ๋ธ์€ simonJJJ์— ์˜ํ•ด ๊ธฐ์—ฌ๋˜์—ˆ์Šต๋‹ˆ๋‹ค.

์‚ฌ์šฉ ์˜ˆ์‹œUsage example

๋‹จ์ผ ๋ฏธ๋””์–ด ์ถ”๋ก Single Media inference

์ด ๋ชจ๋ธ์€ ์ด๋ฏธ์ง€์™€ ๋น„๋””์˜ค๋ฅผ ๋ชจ๋‘ ์ธํ’‹์œผ๋กœ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ๋‹ค์Œ์€ ์ถ”๋ก ์„ ์œ„ํ•œ ์˜ˆ์ œ ์ฝ”๋“œ์ž…๋‹ˆ๋‹ค.


import torch
from transformers import Qwen2VLForConditionalGeneration, AutoTokenizer, AutoProcessor

# ์‚ฌ์šฉ ๊ฐ€๋Šฅํ•œ ์žฅ์น˜์—์„œ ๋ชจ๋ธ์„ ๋ฐ˜ ์ •๋ฐ€๋„(half-precision)๋กœ ๋กœ๋“œ
model = Qwen2VLForConditionalGeneration.from_pretrained("Qwen/Qwen2-VL-7B-Instruct", device_map="auto")
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-VL-7B-Instruct")


conversation = [
    {
        "role":"user",
        "content":[
            {
                "type":"image",
                "url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"
            },
            {
                "type":"text",
                "text":"Describe this image."
            }
        ]
    }
]

inputs = processor.apply_chat_template(
    conversation,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt"
).to(model.device)

# ์ถ”๋ก : ์•„์›ƒํ’‹ ์ƒ์„ฑ
output_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]
output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)
print(output_text)



# ๋น„๋””์˜ค
conversation = [
    {
        "role": "user",
        "content": [
            {"type": "video", "path": "/path/to/video.mp4"},
            {"type": "text", "text": "What happened in the video?"},
        ],
    }
]

inputs = processor.apply_chat_template(
    conversation,
    fps=1,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt"
).to(model.device)


# ์ถ”๋ก : ์•„์›ƒํ’‹ ์ƒ์„ฑ
output_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]
output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)
print(output_text)

๋ฐฐ์น˜ ํ˜ผํ•ฉ ๋ฏธ๋””์–ด ์ถ”๋ก Batch Mixed Media Inference

์ด ๋ชจ๋ธ์€ ์ด๋ฏธ์ง€, ๋น„๋””์˜ค, ํ…์ŠคํŠธ ๋“ฑ ๋‹ค์–‘ํ•œ ์œ ํ˜•์˜ ๋ฐ์ดํ„ฐ๋ฅผ ํ˜ผํ•ฉํ•˜์—ฌ ๋ฐฐ์น˜ ์ž…๋ ฅ์œผ๋กœ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ๋‹ค์Œ์€ ์˜ˆ์ œ์ž…๋‹ˆ๋‹ค.


# ์ฒซ๋ฒˆ์งธ ์ด๋ฏธ์ง€์— ๋Œ€ํ•œ ๋Œ€ํ™”
conversation1 = [
    {
        "role": "user",
        "content": [
            {"type": "image", "path": "/path/to/image1.jpg"},
            {"type": "text", "text": "Describe this image."}
        ]
    }
]

# ๋‘ ๊ฐœ์˜ ์ด๋ฏธ์ง€์— ๋Œ€ํ•œ ๋Œ€ํ™”
conversation2 = [
    {
        "role": "user",
        "content": [
            {"type": "image", "path": "/path/to/image2.jpg"},
            {"type": "image", "path": "/path/to/image3.jpg"},
            {"type": "text", "text": "What is written in the pictures?"}
        ]
    }
]

# ์ˆœ์ˆ˜ ํ…์ŠคํŠธ๋กœ๋งŒ ์ด๋ฃจ์–ด์ง„ ๋Œ€ํ™”
conversation3 = [
    {
        "role": "user",
        "content": "who are you?"
    }
]


# ํ˜ผํ•ฉ๋œ ๋ฏธ๋””์–ด๋กœ ์ด๋ฃจ์–ด์ง„ ๋Œ€ํ™”
conversation4 = [
    {
        "role": "user",
        "content": [
            {"type": "image", "path": "/path/to/image3.jpg"},
            {"type": "image", "path": "/path/to/image4.jpg"},
            {"type": "video", "path": "/path/to/video.jpg"},
            {"type": "text", "text": "What are the common elements in these medias?"},
        ],
    }
]

conversations = [conversation1, conversation2, conversation3, conversation4]
# ๋ฐฐ์น˜ ์ถ”๋ก ์„ ์œ„ํ•œ ์ค€๋น„
ipnuts = processor.apply_chat_template(
    conversations,
    fps=1,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt"
).to(model.device)


# ๋ฐฐ์น˜ ์ถ”๋ก 
output_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]
output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)
print(output_text)

์‚ฌ์šฉ ํŒUsage Tips

์ด๋ฏธ์ง€ ํ•ด์ƒ๋„ ํŠธ๋ ˆ์ด๋“œ์˜คํ”„Image Resolution trade-off

์ด ๋ชจ๋ธ์€ ๋‹ค์–‘ํ•œ ํ•ด์ƒ๋„์˜ ์ž…๋ ฅ์„ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค. ๋””ํดํŠธ๋กœ ์ž…๋ ฅ์— ๋Œ€ํ•ด ๋„ค์ดํ‹ฐ๋ธŒ(native) ํ•ด์ƒ๋„๋ฅผ ์‚ฌ์šฉํ•˜์ง€๋งŒ, ๋” ๋†’์€ ํ•ด์ƒ๋„๋ฅผ ์ ์šฉํ•˜๋ฉด ์„ฑ๋Šฅ์ด ํ–ฅ์ƒ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ์ด๋Š” ๋” ๋งŽ์€ ์—ฐ์‚ฐ ๋น„์šฉ์„ ์ดˆ๋ž˜ํ•ฉ๋‹ˆ๋‹ค. ์‚ฌ์šฉ์ž๋Š” ์ตœ์ ์˜ ์„ค์ •์„ ์œ„ํ•ด ์ตœ์†Œ ๋ฐ ์ตœ๋Œ€ ํ”ฝ์…€ ์ˆ˜๋ฅผ ์กฐ์ •ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

min_pixels = 224*224
max_pixels = 2048*2048
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-VL-7B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels)

์ œํ•œ๋œ GPU RAM์˜ ๊ฒฝ์šฐ, ๋‹ค์Œ๊ณผ ๊ฐ™์ด ํ•ด์ƒ๋„๋ฅผ ์ค„์ผ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค:

min_pixels = 256*28*28
max_pixels = 1024*28*28 
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-VL-7B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels)

์ด๋ ‡๊ฒŒ ํ•˜๋ฉด ๊ฐ ์ด๋ฏธ์ง€๊ฐ€ 256~1024๊ฐœ์˜ ํ† ํฐ์œผ๋กœ ์ธ์ฝ”๋”ฉ๋ฉ๋‹ˆ๋‹ค. ์—ฌ๊ธฐ์„œ 28์€ ๋ชจ๋ธ์ด 14 ํฌ๊ธฐ์˜ ํŒจ์น˜(patch)์™€ 2์˜ ์‹œ๊ฐ„ ํŒจ์น˜(temporal patch size)๋ฅผ ์‚ฌ์šฉํ•˜๊ธฐ ๋•Œ๋ฌธ์— ๋‚˜์˜จ ๊ฐ’์ž…๋‹ˆ๋‹ค (14 ร— 2 = 28).

๋‹ค์ค‘ ์ด๋ฏธ์ง€ ์ธํ’‹Multiple Image Inputs

๊ธฐ๋ณธ์ ์œผ๋กœ ์ด๋ฏธ์ง€์™€ ๋น„๋””์˜ค ์ฝ˜ํ…์ธ ๋Š” ๋Œ€ํ™”์— ์ง์ ‘ ํฌํ•จ๋ฉ๋‹ˆ๋‹ค. ์—ฌ๋Ÿฌ ๊ฐœ์˜ ์ด๋ฏธ์ง€๋ฅผ ์ฒ˜๋ฆฌํ•  ๋•Œ๋Š” ์ด๋ฏธ์ง€ ๋ฐ ๋น„๋””์˜ค์— ๋ผ๋ฒจ์„ ์ถ”๊ฐ€ํ•˜๋ฉด ์ฐธ์กฐํ•˜๊ธฐ๊ฐ€ ๋” ์‰ฌ์›Œ์ง‘๋‹ˆ๋‹ค. ์‚ฌ์šฉ์ž๋Š” ๋‹ค์Œ ์„ค์ •์„ ํ†ตํ•ด ์ด ๋™์ž‘์„ ์ œ์–ดํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค:

conversation = [
    {
        "role": "user",
        "content": [
            {"type": "image"}, 
            {"type": "text", "text": "Hello, how are you?"}
        ]
    },
    {
        "role": "assistant",
        "content": "I'm doing well, thank you for asking. How can I assist you today?"
    },
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Can you describe these images and video?"}, 
            {"type": "image"}, 
            {"type": "image"}, 
            {"type": "video"}, 
            {"type": "text", "text": "These are from my vacation."}
        ]
    },
    {
        "role": "assistant",
        "content": "I'd be happy to describe the images and video for you. Could you please provide more context about your vacation?"
    },
    {
        "role": "user",
        "content": "It was a trip to the mountains. Can you see the details in the images and video?"
    }
]

# ๋””ํดํŠธ:
prompt_without_id = processor.apply_chat_template(conversation, add_generation_prompt=True)
# ์˜ˆ์ƒ ์•„์›ƒํ’‹: '<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|>Hello, how are you?<|im_end|>\n<|im_start|>assistant\nI'm doing well, thank you for asking. How can I assist you today?<|im_end|>\n<|im_start|>user\nCan you describe these images and video?<|vision_start|><|image_pad|><|vision_end|><|vision_start|><|image_pad|><|vision_end|><|vision_start|><|video_pad|><|vision_end|>These are from my vacation.<|im_end|>\n<|im_start|>assistant\nI'd be happy to describe the images and video for you. Could you please provide more context about your vacation?<|im_end|>\n<|im_start|>user\nIt was a trip to the mountains. Can you see the details in the images and video?<|im_end|>\n<|im_start|>assistant\n'


# id ์ถ”๊ฐ€
prompt_with_id = processor.apply_chat_template(conversation, add_generation_prompt=True, add_vision_id=True)
# ์˜ˆ์ƒ ์•„์›ƒํ’‹: '<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n<|im_start|>user\nPicture 1: <|vision_start|><|image_pad|><|vision_end|>Hello, how are you?<|im_end|>\n<|im_start|>assistant\nI'm doing well, thank you for asking. How can I assist you today?<|im_end|>\n<|im_start|>user\nCan you describe these images and video?Picture 2: <|vision_start|><|image_pad|><|vision_end|>Picture 3: <|vision_start|><|image_pad|><|vision_end|>Video 1: <|vision_start|><|video_pad|><|vision_end|>These are from my vacation.<|im_end|>\n<|im_start|>assistant\nI'd be happy to describe the images and video for you. Could you please provide more context about your vacation?<|im_end|>\n<|im_start|>user\nIt was a trip to the mountains. Can you see the details in the images and video?<|im_end|>\n<|im_start|>assistant\n'

๋น ๋ฅธ ์ƒ์„ฑ์„ ์œ„ํ•œ Flash-Attention 2Flash-Attention 2 to speed up generation

์ฒซ๋ฒˆ์งธ๋กœ, Flash Attention 2์˜ ์ตœ์‹  ๋ฒ„์ „์„ ์„ค์น˜ํ•ฉ๋‹ˆ๋‹ค:

pip install -U flash-attn --no-build-isolation

๋˜ํ•œ, Flash-Attention 2๋ฅผ ์ง€์›ํ•˜๋Š” ํ•˜๋“œ์›จ์–ด๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค. ์ž์„ธํ•œ ๋‚ด์šฉ์€ ๊ณต์‹ ๋ฌธ์„œ์ธ flash attention repository์—์„œ ํ™•์ธํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. FlashAttention-2๋Š” ๋ชจ๋ธ์ด torch.float16 ๋˜๋Š” torch.bfloat16 ํ˜•์‹์œผ๋กœ ๋กœ๋“œ๋œ ๊ฒฝ์šฐ์—๋งŒ ์‚ฌ์šฉํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Flash Attention-2๋ฅผ ์‚ฌ์šฉํ•˜์—ฌ ๋ชจ๋ธ์„ ๋กœ๋“œํ•˜๊ณ  ์‹คํ–‰ํ•˜๋ ค๋ฉด, ๋‹ค์Œ๊ณผ ๊ฐ™์ด ๋ชจ๋ธ์„ ๋กœ๋“œํ•  ๋•Œ attn_implementation="flash_attention_2" ์˜ต์…˜์„ ์ถ”๊ฐ€ํ•˜๋ฉด ๋ฉ๋‹ˆ๋‹ค:

from transformers import Qwen2VLForConditionalGeneration

model = Qwen2VLForConditionalGeneration.from_pretrained(
    "Qwen/Qwen2-VL-7B-Instruct", 
    dtype=torch.bfloat16, 
    attn_implementation="flash_attention_2",
)

Qwen2VLConfig

autodoc Qwen2VLConfig

Qwen2VLImageProcessor

autodoc Qwen2VLImageProcessor - preprocess

Qwen2VLImageProcessorFast

autodoc Qwen2VLImageProcessorFast - preprocess

Qwen2VLProcessor

autodoc Qwen2VLProcessor

Qwen2VLModel

autodoc Qwen2VLModel - forward

Qwen2VLForConditionalGeneration

autodoc Qwen2VLForConditionalGeneration - forward