้กน็›ฎๆ–‡ไปถๅคน

ๆ–‡ไปถ
wehub-resource-sync e06fe8e8c6
Secret Leaks / trufflehog (push) Failing after 1s
Build documentation / build (push) Failing after 1s
Build documentation / build_other_lang (push) Failing after 0s
CodeQL Security Analysis / CodeQL Analysis (push) Failing after 0s
PR CI / pr-ci (push) Failing after 1s
Slow tests on important models (on Push - A10) / Get all modified files (push) Failing after 1s
Slow tests on important models (on Push - A10) / Model CI (push) Has been skipped
Self-hosted runner (benchmark) / Benchmark (aws-g5-4xlarge-cache) (push) Has been cancelled
New model PR merged notification / Notify new model (push) Has been cancelled
Update Transformers metadata / build_and_package (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 11:57:37 +08:00

11 KiB

TVP tvp

PyTorch

๊ฐœ์š” overview

Text-Visual Prompting(TVP) ํ”„๋ ˆ์ž„์›Œํฌ๋Š” Yimeng Zhang, Xin Chen, Jinghan Jia, Sijia Liu, Ke Ding์ด ๋ฐœํ‘œํ•œ ๋…ผ๋ฌธ Text-Visual Prompting for Efficient 2D Temporal Video Grounding์—์„œ ์ œ์•ˆ๋˜์—ˆ์Šต๋‹ˆ๋‹ค.

๋…ผ๋ฌธ์˜ ์ดˆ๋ก์€ ๋‹ค์Œ๊ณผ ๊ฐ™์Šต๋‹ˆ๋‹ค:

๋ณธ ๋…ผ๋ฌธ์—์„œ๋Š” ๊ธธ๊ณ , ํŽธ์ง‘๋˜์ง€ ์•Š์€ ๋น„๋””์˜ค์—์„œ ๋ฌธ์žฅ์œผ๋กœ ์„ค๋ช…๋œ ์ˆœ๊ฐ„์˜ ์‹œ์ž‘/์ข…๋ฃŒ ์‹œ์ ์„ ์˜ˆ์ธกํ•˜๋Š” ๊ฒƒ์„ ๋ชฉํ‘œ๋กœ ํ•˜๋Š” Temporal Video Grounding(TVG) ๋ฌธ์ œ๋ฅผ ๋‹ค๋ฃน๋‹ˆ๋‹ค. ์„ธ๋ฐ€ํ•œ 3D ์‹œ๊ฐ์  ํŠน์ง• ๋•๋ถ„์— TVG ๊ธฐ์ˆ ์€ ์ตœ๊ทผ ๋ช‡ ๋…„ ๋™์•ˆ ๋†€๋ผ์šด ๋ฐœ์ „์„ ์ด๋ค˜์Šต๋‹ˆ๋‹ค. ํ•˜์ง€๋งŒ 3D ํ•ฉ์„ฑ๊ณฑ ์‹ ๊ฒฝ๋ง(CNN)์˜ ๋†’์€ ๋ณต์žก์„ฑ์œผ๋กœ ์ธํ•ด ๋ฐ€๋„ ๋†’์€ 3D ์‹œ๊ฐ์  ํŠน์ง•์„ ์ถ”์ถœํ•˜๋Š” ๋ฐ ์‹œ๊ฐ„์ด ์˜ค๋ž˜ ๊ฑธ๋ฆฌ๊ณ  ๊ทธ๋งŒํผ ๋งŽ์€ ๋ฉ”๋ชจ๋ฆฌ์™€ ์—ฐ์‚ฐ ์ž์›์„ ํ•„์š”๋กœ ํ•ฉ๋‹ˆ๋‹ค. ํšจ์œจ์ ์ธ TVG๋ฅผ ์œ„ํ•ด, ๋ณธ ๋…ผ๋ฌธ์—์„œ๋Š” TVG ๋ชจ๋ธ์˜ ์‹œ๊ฐ์  ์ž…๋ ฅ๊ณผ ํ…์ŠคํŠธ ํŠน์ง• ๋ชจ๋‘์— ์ตœ์ ํ™”๋œ ๊ต๋ž€ ํŒจํ„ด('ํ”„๋กฌํ”„ํŠธ'๋ผ๊ณ  ๋ถ€๋ฆ„)์„ ํ†ตํ•ฉํ•˜๋Š” ์ƒˆ๋กœ์šด Text-Visual Prompting(TVP) ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•ฉ๋‹ˆ๋‹ค. 3D CNN๊ณผ ๋šœ๋ ท์ด ๋Œ€๋น„๋˜๊ฒŒ TVP๊ฐ€ 2D TVG ๋ชจ๋ธ์—์„œ ๋น„์ „ ์ธ์ฝ”๋”์™€ ์–ธ์–ด ์ธ์ฝ”๋”๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ๊ณต๋™ ํ•™์Šตํ•  ์ˆ˜ ์žˆ๊ฒŒ ํ•˜๊ณ , ๋‚ฎ์€ ๋ณต์žก๋„์˜ ํฌ์†Œํ•œ 2D ์‹œ๊ฐ์  ํŠน์ง•๋งŒ์„ ์‚ฌ์šฉํ•˜์—ฌ ํฌ๋กœ์Šค ๋ชจ๋‹ฌ ํŠน์ง• ์œตํ•ฉ์˜ ์„ฑ๋Šฅ์„ ํ–ฅ์ƒ์‹œํ‚ต๋‹ˆ๋‹ค. ๋” ๋‚˜์•„๊ฐ€, TVG์˜ ํšจ์œจ์ ์ธ ํ•™์Šต์„ ์œ„ํ•ด Temporal-Distance IoU(TDIoU) ์†์‹ค ํ•จ์ˆ˜๋ฅผ ์ œ์•ˆํ•ฉ๋‹ˆ๋‹ค. ๋‘ ๊ฐœ์˜ ๋ฒค์น˜๋งˆํฌ ๋ฐ์ดํ„ฐ ์„ธํŠธ์ธ Charades-STA์™€ ActivityNet Captions ๋ฐ์ดํ„ฐ์…‹์— ๋Œ€ํ•œ ์‹คํ—˜์„ ํ†ตํ•ด, ์ œ์•ˆ๋œ TVP๊ฐ€ 2D TVG์˜ ์„ฑ๋Šฅ์„ ํฌ๊ฒŒ ํ–ฅ์ƒ์‹œํ‚ค๊ณ (์˜ˆ: Charades-STA์—์„œ 9.79% ํ–ฅ์ƒ, ActivityNet Captions์—์„œ 30.77% ํ–ฅ์ƒ) 3D ์‹œ๊ฐ์  ํŠน์ง•์„ ์‚ฌ์šฉํ•˜๋Š” TVG์— ๋น„ํ•ด 5๋ฐฐ์˜ ์ถ”๋ก  ๊ฐ€์†์„ ๋‹ฌ์„ฑํ•จ์„ ์‹คํ—˜์ ์œผ๋กœ ์ž…์ฆํ•ฉ๋‹ˆ๋‹ค.

์ด ์—ฐ๊ตฌ๋Š” Temporal Video Grounding(TVG)์„ ๋‹ค๋ฃน๋‹ˆ๋‹ค. TVG๋Š” ๋ฌธ์žฅ์œผ๋กœ ์„ค๋ช…๋œ ํŠน์ • ์ด๋ฒคํŠธ์˜ ์‹œ์ž‘ ๋ฐ ์ข…๋ฃŒ ์‹œ์ ์„ ๊ธด ๋น„๋””์˜ค์—์„œ ์ •ํ™•ํžˆ ์ฐพ์•„๋‚ด๋Š” ๊ณผ์ •์ž…๋‹ˆ๋‹ค. TVG ์„ฑ๋Šฅ์„ ํ–ฅ์ƒ์‹œํ‚ค๊ธฐ ์œ„ํ•ด Text-Visual Prompting(TVP)์ด ์ œ์•ˆ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. TVP๋Š” 'ํ”„๋กฌํ”„ํŠธ'๋ผ๊ณ  ์•Œ๋ ค์ง„ ํŠน๋ณ„ํžˆ ์„ค๊ณ„๋œ ํŒจํ„ด์„ TVG ๋ชจ๋ธ์˜ ์‹œ๊ฐ์ (์ด๋ฏธ์ง€ ๊ธฐ๋ฐ˜) ๋ฐ ํ…์ŠคํŠธ(๋‹จ์–ด ๊ธฐ๋ฐ˜) ์ž…๋ ฅ ๊ตฌ์„ฑ ์š”์†Œ ๋ชจ๋‘์— ํ†ตํ•ฉํ•˜๋Š” ๊ฒƒ์„ ๋ฐฉ์‹์ž…๋‹ˆ๋‹ค. ์ด ํ”„๋กฌํ”„ํŠธ๋Š” ์ถ”๊ฐ€์ ์ธ ์‹œ๊ณต๊ฐ„์  ์ปจํ…์ŠคํŠธ๋ฅผ ์ œ๊ณตํ•จ์œผ๋กœ์จ ๋ชจ๋ธ์ด ๋น„๋””์˜ค ๋‚ด ์ด๋ฒคํŠธ ์‹œ์ ์˜ ์˜ˆ์ธก ์ •ํ™•๋„๋ฅผ ๋†’์ž…๋‹ˆ๋‹ค. ์ด ์ ‘๊ทผ ๋ฐฉ์‹์€ 3D ์‹œ๊ฐ์  ์ž…๋ ฅ ๋Œ€์‹  2D ์ž…๋ ฅ์„ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค. 3D ์ž…๋ ฅ์€ ๋ณด๋‹ค ํ’๋ถ€ํ•œ ์‹œ๊ณต๊ฐ„์  ์„ธ๋ถ€ ์ •๋ณด๋ฅผ ์ œ๊ณตํ•˜์ง€๋งŒ ์ฒ˜๋ฆฌํ•˜๋Š” ๋ฐ ์‹œ๊ฐ„์ด ๋” ๋งŽ์ด ๊ฑธ๋ฆฝ๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ ํ”„๋กฌํ”„ํŒ… ๋ฉ”์†Œ๋“œ์™€ ํ•จ๊ป˜ 2D ์ž…๋ ฅ์„ ์‚ฌ์šฉํ•˜์—ฌ ์ด์™€ ์œ ์‚ฌํ•œ ์ˆ˜์ค€์˜ ์ปจํ…์ŠคํŠธ์™€ ์ •ํ™•๋„๋ฅผ ๋” ํšจ์œจ์ ์œผ๋กœ ์ œ๊ณตํ•˜๋Š” ๊ฒƒ์„ ๋ชฉํ‘œ๋กœ ํ•ฉ๋‹ˆ๋‹ค.

drawing

TVP ์•„ํ‚คํ…์ฒ˜. ์›๋ณธ ๋…ผ๋ฌธ์—์„œ ๋ฐœ์ทŒ.

์ด ๋ชจ๋ธ์€ Jiqing Feng๋‹˜์ด ๊ธฐ์—ฌํ–ˆ์Šต๋‹ˆ๋‹ค. ์›๋ณธ ์ฝ”๋“œ๋Š” ์ด ๊ณณ์—์„œ ์ฐพ์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

์‚ฌ์šฉ ํŒ ๋ฐ ์˜ˆ์‹œ usage-tips-and-examples

ํ”„๋กฌํ”„ํŠธ๋Š” ์ตœ์ ํ™”๋œ ๊ต๋ž€ ํŒจํ„ด์œผ๋กœ ์ž…๋ ฅ ๋น„๋””์˜ค ํ”„๋ ˆ์ž„์ด๋‚˜ ํ…์ŠคํŠธ ํŠน์ง•์— ์ถ”๊ฐ€๋˜๋Š” ํŒจํ„ด์ž…๋‹ˆ๋‹ค. ๋ฒ”์šฉ ์„ธํŠธ๋ž€ ๋ชจ๋“  ์ž…๋ ฅ์— ๋Œ€ํ•ด ๋™์ผํ•œ ํ”„๋กฌํ”„ํŠธ ์„ธํŠธ๋ฅผ ์‚ฌ์šฉํ•˜๋Š” ๊ฒƒ์„ ๋งํ•ฉ๋‹ˆ๋‹ค. ์ฆ‰, ์ž…๋ ฅ ๋‚ด์šฉ๊ณผ ๊ด€๊ณ„์—†์ด ๋ชจ๋“  ๋น„๋””์˜ค ํ”„๋ ˆ์ž„๊ณผ ํ…์ŠคํŠธ ํŠน์ง•์— ์ด ํ”„๋กฌํ”„ํŠธ๋“ค์„ ์ผ๊ด€์ ์œผ๋กœ ์ถ”๊ฐ€ํ•ฉ๋‹ˆ๋‹ค.

TVP๋Š” ์‹œ๊ฐ ์ธ์ฝ”๋”์™€ ํฌ๋กœ์Šค ๋ชจ๋‹ฌ ์ธ์ฝ”๋”๋กœ ๊ตฌ์„ฑ๋ฉ๋‹ˆ๋‹ค. ๋ฒ”์šฉ ์‹œ๊ฐ ํ”„๋กฌํ”„ํŠธ์™€ ํ…์ŠคํŠธ ํ”„๋กฌํ”„ํŠธ ์„ธํŠธ๊ฐ€ ๊ฐ๊ฐ ์ƒ˜ํ”Œ๋ง๋œ ๋น„๋””์˜ค ํ”„๋ ˆ์ž„๊ณผ ํ…์ŠคํŠธ ํŠน์ง•์— ํ†ตํ•ฉ๋ฉ๋‹ˆ๋‹ค. ํŠนํžˆ, ์„œ๋กœ ๋‹ค๋ฅธ ์‹œ๊ฐ ํ”„๋กฌํ”„ํŠธ ์„ธํŠธ๊ฐ€ ํŽธ์ง‘๋˜์ง€ ์•Š์€ ํ•œ ๋น„๋””์˜ค์—์„œ ๊ท ์ผํ•˜๊ฒŒ ์ƒ˜ํ”Œ๋ง๋œ ํ”„๋ ˆ์ž„์— ์ˆœ์„œ๋Œ€๋กœ ์ ์šฉ๋ฉ๋‹ˆ๋‹ค.

์ด ๋ชจ๋ธ์˜ ๋ชฉํ‘œ๋Š” ํ•™์Šต ๊ฐ€๋Šฅํ•œ ํ”„๋กฌํ”„ํŠธ๋ฅผ ์‹œ๊ฐ์  ์ž…๋ ฅ๊ณผ ํ…์ŠคํŠธ ํŠน์ง• ๋ชจ๋‘์— ํ†ตํ•ฉํ•˜์—ฌ Temporal Video Grounding(TVG) ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๋Š” ๊ฒƒ์ž…๋‹ˆ๋‹ค.

์›์น™์ ์œผ๋กœ, ์ œ์•ˆ๋œ ์•„ํ‚คํ…์ฒ˜์—๋Š” ์–ด๋–ค ์‹œ๊ฐ ์ธ์ฝ”๋”๋‚˜ ํฌ๋กœ์Šค ๋ชจ๋‹ฌ ์ธ์ฝ”๋”๋ผ๋„ ์ ์šฉํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

[TvpProcessor]๋Š” [BertTokenizer]์™€ [TvpImageProcessor]๋ฅผ ๋‹จ์ผ ์ธ์Šคํ„ด์Šค๋กœ ๋ž˜ํ•‘ํ•˜์—ฌ ํ…์ŠคํŠธ๋ฅผ ์ธ์ฝ”๋”ฉํ•˜๊ณ  ์ด๋ฏธ์ง€๋ฅผ ๊ฐ๊ฐ ์ค€๋น„ํ•ฉ๋‹ˆ๋‹ค.

๋‹ค์Œ ์˜ˆ์‹œ๋Š” [TvpProcessor]์™€ [TvpForVideoGrounding]์„ ์‚ฌ์šฉํ•˜์—ฌ TVG๋ฅผ ์‹คํ–‰ํ•˜๋Š” ๋ฐฉ๋ฒ•์„ ๋ณด์—ฌ์ค๋‹ˆ๋‹ค.

import av
import cv2
import numpy as np
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoProcessor, TvpForVideoGrounding


def pyav_decode(container, sampling_rate, num_frames, clip_idx, num_clips, target_fps):
    '''
    ์›๋ณธ fps์˜ ๋น„๋””์˜ค๋ฅผ ์ง€์ •ํ•œ fps(target_fps)๋กœ ๋ณ€ํ™˜ํ•˜๊ณ  PyAV ๋””์ฝ”๋”๋กœ ๋น„๋””์˜ค๋ฅผ ๋””์ฝ”๋”ฉํ•ฉ๋‹ˆ๋‹ค.
    Args:
        container (container): pyav ์ปจํ…Œ์ด๋„ˆ ๊ฐ์ฒด์ž…๋‹ˆ๋‹ค.
        sampling_rate (int): ํ”„๋ ˆ์ž„ ์ƒ˜ํ”Œ๋ง ์†๋„์ž…๋‹ˆ๋‹ค.(์ƒ˜ํ”Œ๋ง๋œ ๋‘๊ฐœ์˜ ํ”„๋ ˆ์ž„ ์‚ฌ์ด์˜ ๊ฐ„๊ฒฉ์„ ๋งํ•ฉ๋‹ˆ๋‹ค)
        num_frames (int): ์ƒ˜ํ”Œ๋งํ•  ํ”„๋ ˆ์ž„ ์ˆ˜์ž…๋‹ˆ๋‹ค.
        clip_idx (int): clip_idx๊ฐ€ -1์ด๋ฉด ์‹œ๊ฐ„ ์ถ•์—์„œ ๋ฌด์ž‘์œ„ ์ƒ˜ํ”Œ๋ง์„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค.
            clip_idx๊ฐ€ -1๋ณด๋‹ค ํฌ๋ฉด ๋น„๋””์˜ค๋ฅผ num_clips ๊ฐœ๋กœ ๊ท ๋“ฑ ๋ถ„ํ• ํ•œ ํ›„
            clip_idx๋ฒˆ์งธ ๋น„๋””์˜ค ํด๋ฆฝ์„ ์„ ํƒํ•ฉ๋‹ˆ๋‹ค.
        num_clips (int): ์ฃผ์–ด์ง„ ๋น„๋””์˜ค์—์„œ ๊ท ์ผํ•˜๊ฒŒ ์ƒ˜ํ”Œ๋งํ•  ์ „์ฒด ํด๋ฆฝ ์ˆ˜์ž…๋‹ˆ๋‹ค.
        target_fps (int): ์ž…๋ ฅ ๋น„๋””์˜ค์˜ fps๊ฐ€ ๋‹ค๋ฅผ ์ˆ˜ ์žˆ์œผ๋ฏ€๋กœ, ์ƒ˜ํ”Œ๋ง ์ „์—
            ์ง€์ •ํ•œ fps๋กœ ๋ณ€ํ™˜ํ•ฉ๋‹ˆ๋‹ค
    Returns:
        frames (tensor): ๋น„๋””์˜ค์—์„œ ๋””์ฝ”๋”ฉ๋œ ํ”„๋ ˆ์ž„์ž…๋‹ˆ๋‹ค. ๋น„๋””์˜ค ์ŠคํŠธ๋ฆผ์„ ์ฐพ์„ ์ˆ˜ ์—†๋Š” ๊ฒฝ์šฐ
            None์„ ๋ฐ˜ํ™˜ํ•ฉ๋‹ˆ๋‹ค.
        fps (float): ๋น„๋””์˜ค์˜ ์ดˆ๋‹น ํ”„๋ ˆ์ž„ ์ˆ˜์ž…๋‹ˆ๋‹ค.
    '''
    video = container.streams.video[0]
    fps = float(video.average_rate)
    clip_size = sampling_rate * num_frames / target_fps * fps
    delta = max(num_frames - clip_size, 0)
    start_idx = delta * clip_idx / num_clips
    end_idx = start_idx + clip_size - 1
    timebase = video.duration / num_frames
    video_start_pts = int(start_idx * timebase)
    video_end_pts = int(end_idx * timebase)
    seek_offset = max(video_start_pts - 1024, 0)
    container.seek(seek_offset, any_frame=False, backward=True, stream=video)
    frames = {}
    for frame in container.decode(video=0):
        if frame.pts < video_start_pts:
            continue
        frames[frame.pts] = frame
        if frame.pts > video_end_pts:
            break
    frames = [frames[pts] for pts in sorted(frames)]
    return frames, fps


def decode(container, sampling_rate, num_frames, clip_idx, num_clips, target_fps):
    '''
    ๋น„๋””์˜ค๋ฅผ ๋””์ฝ”๋”ฉํ•˜๊ณ  ์‹œ๊ฐ„ ์ถ• ์ƒ˜ํ”Œ๋ง์„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค.
    Args:
        container (container): pyav ์ปจํ…Œ์ด๋„ˆ ๊ฐ์ฒด์ž…๋‹ˆ๋‹ค.
        sampling_rate (int): ํ”„๋ ˆ์ž„ ์ƒ˜ํ”Œ๋ง ์†๋„์ž…๋‹ˆ๋‹ค.(์ƒ˜ํ”Œ๋ง๋œ ๋‘๊ฐœ์˜ ํ”„๋ ˆ์ž„ ์‚ฌ์ด์˜ ๊ฐ„๊ฒฉ์„ ๋งํ•ฉ๋‹ˆ๋‹ค)
        num_frames (int): ์ƒ˜ํ”Œ๋งํ•  ํ”„๋ ˆ์ž„ ์ˆ˜์ž…๋‹ˆ๋‹ค.
        clip_idx (int): clip_idx๊ฐ€ -1์ด๋ฉด ์‹œ๊ฐ„ ์ถ•์—์„œ ๋ฌด์ž‘์œ„ ์ƒ˜ํ”Œ๋ง์„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค.
            clip_idx๊ฐ€ -1๋ณด๋‹ค ํฌ๋ฉด ๋น„๋””์˜ค๋ฅผ num_clips ๊ฐœ๋กœ ๊ท ๋“ฑ ๋ถ„ํ• ํ•œ ํ›„
            clip_idx๋ฒˆ์งธ ๋น„๋””์˜ค ํด๋ฆฝ์„ ์„ ํƒํ•ฉ๋‹ˆ๋‹ค.
        num_clips (int): ์ฃผ์–ด์ง„ ๋น„๋””์˜ค์—์„œ ๊ท ์ผํ•˜๊ฒŒ ์ƒ˜ํ”Œ๋งํ•  ์ „์ฒด ํด๋ฆฝ ์ˆ˜์ž…๋‹ˆ๋‹ค.
        target_fps (int): ์ž…๋ ฅ ๋น„๋””์˜ค์˜ fps๊ฐ€ ๋‹ค๋ฅผ ์ˆ˜ ์žˆ์œผ๋ฏ€๋กœ, ์ƒ˜ํ”Œ๋ง ์ „์—
            ์ง€์ •ํ•œ fps๋กœ ๋ณ€ํ™˜ํ•ฉ๋‹ˆ๋‹ค
    Returns:
        frames (tensor): ๋น„๋””์˜ค์—์„œ ๋””์ฝ”๋”ฉ๋œ ํ”„๋ ˆ์ž„์ž…๋‹ˆ๋‹ค.
    '''
    assert clip_idx >= -2, "Not a valid clip_idx {}".format(clip_idx)
    frames, fps = pyav_decode(container, sampling_rate, num_frames, clip_idx, num_clips, target_fps)
    clip_size = sampling_rate * num_frames / target_fps * fps
    index = np.linspace(0, clip_size - 1, num_frames)
    index = np.clip(index, 0, len(frames) - 1).astype(np.int64)
    frames = np.array([frames[idx].to_rgb().to_ndarray() for idx in index])
    frames = frames.transpose(0, 3, 1, 2)
    return frames


file = hf_hub_download(repo_id="Intel/tvp_demo", filename="AK2KG.mp4", repo_type="dataset")
model = TvpForVideoGrounding.from_pretrained("Intel/tvp-base")

decoder_kwargs = dict(
    container=av.open(file, metadata_errors="ignore"),
    sampling_rate=1,
    num_frames=model.config.num_frames,
    clip_idx=0,
    num_clips=1,
    target_fps=3,
)
raw_sampled_frms = decode(**decoder_kwargs)

text = "a person is sitting on a bed."
processor = AutoProcessor.from_pretrained("Intel/tvp-base")
model_inputs = processor(
    text=[text], videos=list(raw_sampled_frms), return_tensors="pt", max_text_length=100#, size=size
)

model_inputs["pixel_values"] = model_inputs["pixel_values"].to(model.dtype)
output = model(**model_inputs)

def get_video_duration(filename):
    cap = cv2.VideoCapture(filename)
    if cap.isOpened():
        rate = cap.get(5)
        frame_num = cap.get(7)
        duration = frame_num/rate
        return duration
    return -1

duration = get_video_duration(file)
start, end = processor.post_process_video_grounding(output.logits, duration)

print(f"The time slot of the video corresponding to the text \"{text}\" is from {start}s to {end}s")

ํŒ:

  • ์ด TVP ๊ตฌํ˜„์€ ํ…์ŠคํŠธ ์ž„๋ฒ ๋”ฉ์„ ์ƒ์„ฑํ•˜๊ธฐ ์œ„ํ•ด [BertTokenizer]๋ฅผ ์‚ฌ์šฉํ•˜๊ณ , ์‹œ๊ฐ์  ์ž„๋ฒ ๋”ฉ์„ ๊ณ„์‚ฐํ•˜๊ธฐ ์œ„ํ•ด Resnet-50 ๋ชจ๋ธ์„ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค.
  • ์‚ฌ์ „ ํ•™์Šต๋œ tvp-base์˜ ์ฒดํฌํฌ์ธํŠธ๊ฐ€ ๊ณต๊ฐœ๋˜์–ด ์žˆ์Šต๋‹ˆ๋‹ค.
  • ์‹œ๊ฐ„์  ๋น„๋””์˜ค ๊ทธ๋ผ์šด๋”ฉ ์ž‘์—…์— ๋Œ€ํ•œ TVP์˜ ์„ฑ๋Šฅ์€ ํ‘œ 2๋ฅผ ์ฐธ๊ณ ํ•˜์„ธ์š”.

TvpConfig transformers.TvpConfig

autodoc TvpConfig

TvpImageProcessor transformers.TvpImageProcessor

autodoc TvpImageProcessor - preprocess

TvpProcessor transformers.TvpProcessor

autodoc TvpProcessor - call

TvpModel transformers.TvpModel

autodoc TvpModel - forward

TvpForVideoGrounding transformers.TvpForVideoGrounding

autodoc TvpForVideoGrounding - forward