้กน็›ฎๆ–‡ไปถๅคน

ๆ–‡ไปถ
wehub-resource-sync e06fe8e8c6
Secret Leaks / trufflehog (push) Failing after 1s
Build documentation / build (push) Failing after 1s
Build documentation / build_other_lang (push) Failing after 0s
CodeQL Security Analysis / CodeQL Analysis (push) Failing after 0s
PR CI / pr-ci (push) Failing after 1s
Slow tests on important models (on Push - A10) / Get all modified files (push) Failing after 1s
Slow tests on important models (on Push - A10) / Model CI (push) Has been skipped
Self-hosted runner (benchmark) / Benchmark (aws-g5-4xlarge-cache) (push) Has been cancelled
New model PR merged notification / Notify new model (push) Has been cancelled
Update Transformers metadata / build_and_package (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 11:57:37 +08:00

13 KiB

๋„๊ตฌ์™€ RAGTools-and-RAG

[~PreTrainedTokenizerBase.apply_chat_template] ๋ฉ”์†Œ๋“œ๋Š” ์ฑ„ํŒ… ๋ฉ”์‹œ์ง€ ์™ธ์—๋„ ๋ฌธ์ž์—ด, ๋ฆฌ์ŠคํŠธ, ๋”•์…”๋„ˆ๋ฆฌ ๋“ฑ ๊ฑฐ์˜ ๋ชจ๋“  ์ข…๋ฅ˜์˜ ์ถ”๊ฐ€ ์ธ์ˆ˜ ํƒ€์ž…์„ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค. ์ด๋ฅผ ํ†ตํ•ด ๋‹ค์–‘ํ•œ ์‚ฌ์šฉ ์ƒํ™ฉ์—์„œ ์ฑ„ํŒ… ํ…œํ”Œ๋ฆฟ์„ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

์ด ๊ฐ€์ด๋“œ์—์„œ๋Š” ๋„๊ตฌ ๋ฐ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG)๊ณผ ํ•จ๊ป˜ ์ฑ„ํŒ… ํ…œํ”Œ๋ฆฟ์„ ์‚ฌ์šฉํ•˜๋Š” ๋ฐฉ๋ฒ•์„ ๋ณด์—ฌ๋“œ๋ฆฝ๋‹ˆ๋‹ค.

๋„๊ตฌTools

๋„๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ํŠน์ • ์ž‘์—…์„ ์ˆ˜ํ–‰ํ•˜๊ธฐ ์œ„ํ•ด ํ˜ธ์ถœํ•  ์ˆ˜ ์žˆ๋Š” ํ•จ์ˆ˜์ž…๋‹ˆ๋‹ค. ์ด๋Š” ์‹ค์‹œ๊ฐ„ ์ •๋ณด, ๊ณ„์‚ฐ ๋„๊ตฌ ๋˜๋Š” ๋Œ€๊ทœ๋ชจ ๋ฐ์ดํ„ฐ๋ฒ ์ด์Šค ์ ‘๊ทผ ๋“ฑ์„ ํ†ตํ•ด ๋Œ€ํ™”ํ˜• ์—์ด์ „ํŠธ์˜ ๊ธฐ๋Šฅ์„ ํ™•์žฅํ•˜๋Š” ๊ฐ•๋ ฅํ•œ ๋ฐฉ๋ฒ•์ž…๋‹ˆ๋‹ค.

๋„๊ตฌ๋ฅผ ๋งŒ๋“ค ๋•Œ๋Š” ์•„๋ž˜ ๊ทœ์น™์„ ๋”ฐ๋ฅด์„ธ์š”.

  1. ํ•จ์ˆ˜๋Š” ๊ธฐ๋Šฅ์„ ์ž˜ ์„ค๋ช…ํ•˜๋Š” ์ด๋ฆ„์„ ๊ฐ€์ ธ์•ผ ํ•ฉ๋‹ˆ๋‹ค.
  2. ํ•จ์ˆ˜์˜ ์ธ์ˆ˜๋Š” ํ•จ์ˆ˜ ํ—ค๋”์— ํƒ€์ž… ํžŒํŠธ๋ฅผ ํฌํ•จํ•ด์•ผ ํ•ฉ๋‹ˆ๋‹ค(Args ๋ธ”๋ก์—๋Š” ํฌํ•จํ•˜์ง€ ๋งˆ์„ธ์š”).
  3. ํ•จ์ˆ˜์—๋Š” Google ์Šคํƒ€์ผ ์˜ ๋…์ŠคํŠธ๋ง(docstring)์ด ํฌํ•จ๋˜์–ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.
  4. ํ•จ์ˆ˜์— ๋ฐ˜ํ™˜ ํƒ€์ž…๊ณผ Returns ๋ธ”๋ก์„ ํฌํ•จํ•  ์ˆ˜ ์žˆ์ง€๋งŒ, ๋„๊ตฌ๋ฅผ ํ™œ์šฉํ•˜๋Š” ๋Œ€๋ถ€๋ถ„์˜ ๋ชจ๋ธ์—์„œ ์ด๋ฅผ ์‚ฌ์šฉํ•˜์ง€ ์•Š๊ธฐ ๋•Œ๋ฌธ์— ๋ฌด์‹œํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

์ฃผ์–ด์ง„ ์œ„์น˜์˜ ํ˜„์žฌ ์˜จ๋„์™€ ํ’์†์„ ๊ฐ€์ ธ์˜ค๋Š” ๋„๊ตฌ์˜ ์˜ˆ์‹œ๋Š” ์•„๋ž˜์™€ ๊ฐ™์Šต๋‹ˆ๋‹ค.

def get_current_temperature(location: str, unit: str) -> float:
    """
    ์ฃผ์–ด์ง„ ์œ„์น˜์˜ ํ˜„์žฌ ์˜จ๋„๋ฅผ ๊ฐ€์ ธ์˜ต๋‹ˆ๋‹ค.
    
    Args:
        location: ์˜จ๋„๋ฅผ ๊ฐ€์ ธ์˜ฌ ์œ„์น˜, "๋„์‹œ, ๊ตญ๊ฐ€" ํ˜•์‹
        unit: ์˜จ๋„๋ฅผ ๋ฐ˜ํ™˜ํ•  ๋‹จ์œ„. (์„ ํƒ์ง€: ["celsius(์„ญ์”จ)", "fahrenheit(ํ™”์”จ)"])
    Returns:
        ์ฃผ์–ด์ง„ ์œ„์น˜์˜ ์ง€์ •๋œ ๋‹จ์œ„๋กœ ํ‘œ์‹œ๋œ ํ˜„์žฌ ์˜จ๋„(float ์ž๋ฃŒํ˜•).
    """
    return 22.  # ์‹ค์ œ ํ•จ์ˆ˜๋ผ๋ฉด ์•„๋งˆ ์ง„์งœ๋กœ ๊ธฐ์˜จ์„ ๊ฐ€์ ธ์™€์•ผ๊ฒ ์ฃ !

def get_current_wind_speed(location: str) -> float:
    """
    ์ฃผ์–ด์ง„ ์œ„์น˜์˜ ํ˜„์žฌ ํ’์†์„ km/h ๋‹จ์œ„๋กœ ๊ฐ€์ ธ์˜ต๋‹ˆ๋‹ค.
    
    Args:
        location: ์˜จ๋„๋ฅผ ๊ฐ€์ ธ์˜ฌ ์œ„์น˜, "๋„์‹œ, ๊ตญ๊ฐ€" ํ˜•์‹
    Returns:
        ์ฃผ์–ด์ง„ ์œ„์น˜์˜ ํ˜„์žฌ ํ’์†(km/h, float ์ž๋ฃŒํ˜•).
    """
    return 6.  # ์‹ค์ œ ํ•จ์ˆ˜๋ผ๋ฉด ์•„๋งˆ ์ง„์งœ๋กœ ํ’์†์„ ๊ฐ€์ ธ์™€์•ผ๊ฒ ์ฃ !

tools = [get_current_temperature, get_current_wind_speed]

NousResearch/Hermes-2-Pro-Llama-3-8B์™€ ๊ฐ™์ด ๋„๊ตฌ ์‚ฌ์šฉ์„ ์ง€์›ํ•˜๋Š” ๋ชจ๋ธ๊ณผ ํ† ํฌ๋‚˜์ด์ €๋ฅผ ๊ฐ€์ ธ์˜ค์„ธ์š”. ํ•˜๋“œ์›จ์–ด๊ฐ€ ์ง€์›๋œ๋‹ค๋ฉด Command-R์ด๋‚˜ Mixtral-8x22B์™€ ๊ฐ™์€ ๋” ํฐ ๋ชจ๋ธ๋„ ๊ณ ๋ คํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained( "NousResearch/Hermes-2-Pro-Llama-3-8B")
tokenizer = AutoTokenizer.from_pretrained( "NousResearch/Hermes-2-Pro-Llama-3-8B")
model = AutoModelForCausalLM.from_pretrained( "NousResearch/Hermes-2-Pro-Llama-3-8B", torch_dtype=torch.bfloat16, device_map="auto")

์ฑ„ํŒ… ๋ฉ”์‹œ์ง€๋ฅผ ์ƒ์„ฑํ•ฉ๋‹ˆ๋‹ค.

messages = [
  {"role": "system", "content": "You are a bot that responds to weather queries. You should reply with the unit used in the queried location."},
  {"role": "user", "content": "Hey, what's the temperature in Paris right now?"}
]

messages์™€ ๋„๊ตฌ ๋ชฉ๋ก tools๋ฅผ [~PreTrainedTokenizerBase.apply_chat_template]์— ์ „๋‹ฌํ•œ ๋’ค, ์ด๋ฅผ ๋ชจ๋ธ์˜ ์ž…๋ ฅ์œผ๋กœ ์‚ฌ์šฉํ•˜์—ฌ ํ…์ŠคํŠธ๋ฅผ ์ƒ์„ฑํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

inputs = tokenizer.apply_chat_template(messages, tools=tools, add_generation_prompt=True, return_dict=True, return_tensors="pt")
inputs = {k: v for k, v in inputs.items()}
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0][len(inputs["input_ids"][0]):]))
<tool_call>
{"arguments": {"location": "Paris, France", "unit": "celsius"}, "name": "get_current_temperature"}
</tool_call><|im_end|>

์ฑ„ํŒ… ๋ชจ๋ธ์€ ๋…์ŠคํŠธ๋ง(docstring)์— ์ •์˜๋œ ํ˜•์‹์— ๋”ฐ๋ผ get_current_temperature ํ•จ์ˆ˜์— ์˜ฌ๋ฐ”๋ฅธ ๋งค๊ฐœ๋ณ€์ˆ˜๋ฅผ ์ „๋‹ฌํ•ด ํ˜ธ์ถœํ–ˆ์Šต๋‹ˆ๋‹ค. ํŒŒ๋ฆฌ๋ฅผ ๊ธฐ์ค€์œผ๋กœ ์œ„์น˜๋ฅผ ํ”„๋ž‘์Šค๋กœ ์ถ”๋ก ํ–ˆ์œผ๋ฉฐ, ์˜จ๋„ ๋‹จ์œ„๋Š” ์„ญ์”จ๋ฅผ ์‚ฌ์šฉํ•ด์•ผ ํ•œ๋‹ค๊ณ  ํŒ๋‹จํ–ˆ์Šต๋‹ˆ๋‹ค.

์ด์ œ get_current_temperature ํ•จ์ˆ˜์™€ ํ•ด๋‹น ์ธ์ˆ˜๋“ค์„ tool_call ๋”•์…”๋„ˆ๋ฆฌ์— ๋‹ด์•„ ์ฑ„ํŒ… ๋ฉ”์‹œ์ง€์— ์ถ”๊ฐ€ํ•ฉ๋‹ˆ๋‹ค. tool_call ๋”•์…”๋„ˆ๋ฆฌ๋Š” system์ด๋‚˜ user๊ฐ€ ์•„๋‹Œ assistant ์—ญํ• ๋กœ ์ œ๊ณต๋˜์–ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.

Warning

OpenAI API๋Š” tool_call ํ˜•์‹์œผ๋กœ JSON ๋ฌธ์ž์—ด์„ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค. Transformers์—์„œ ์‚ฌ์šฉํ•  ๊ฒฝ์šฐ ๋”•์…”๋„ˆ๋ฆฌ๋ฅผ ์š”๊ตฌํ•˜๊ธฐ ๋•Œ๋ฌธ์—, ์˜ค๋ฅ˜๊ฐ€ ๋ฐœ์ƒํ•˜๊ฑฐ๋‚˜ ๋ชจ๋ธ์ด ์ด์ƒํ•˜๊ฒŒ ๋™์ž‘ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

tool_call = {"name": "get_current_temperature", "arguments": {"location": "Paris, France", "unit": "celsius"}}
messages.append({"role": "assistant", "tool_calls": [{"type": "function", "function": tool_call}]})

์–ด์‹œ์Šคํ„ดํŠธ๊ฐ€ ํ•จ์ˆ˜ ์ถœ๋ ฅ์„ ์ฝ๊ณ  ์‚ฌ์šฉ์ž์™€ ์ฑ„ํŒ…ํ•  ์ˆ˜ ์žˆ๋„๋ก ํ•ฉ๋‹ˆ๋‹ค.

inputs = tokenizer.apply_chat_template(messages, tools=tools, add_generation_prompt=True, return_dict=True, return_tensors="pt")
inputs = {k: v for k, v in inputs.items()}
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][len(inputs["input_ids"][0]):]))
The temperature in Paris, France right now is approximately 12ยฐC (53.6ยฐF).<|im_end|>

Mistral ๋ฐ Mixtral ๋ชจ๋ธ์˜ ๊ฒฝ์šฐ ์ถ”๊ฐ€์ ์œผ๋กœ tool_call_id๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค. tool_call_id๋Š” 9์ž๋ฆฌ ์˜์ˆซ์ž ๋ฌธ์ž์—ด๋กœ ์ƒ์„ฑ๋˜์–ด tool_call ๋”•์…”๋„ˆ๋ฆฌ์˜ id ํ‚ค์— ํ• ๋‹น๋ฉ๋‹ˆ๋‹ค.

tool_call_id = "9Ae3bDc2F"
tool_call = {"name": "get_current_temperature", "arguments": {"location": "Paris, France", "unit": "celsius"}}
messages.append({"role": "assistant", "tool_calls": [{"type": "function", "id": tool_call_id, "function": tool_call}]})
inputs = tokenizer.apply_chat_template(messages, tools=tools, add_generation_prompt=True, return_dict=True, return_tensors="pt")
inputs = {k: v for k, v in inputs.items()}
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][len(inputs["input_ids"][0]):]))

์Šคํ‚ค๋งˆSchema

[~PreTrainedTokenizerBase.apply_chat_template]์€ ํ•จ์ˆ˜๋ฅผ JSON ์Šคํ‚ค๋งˆ๋กœ ๋ณ€ํ™˜ํ•˜์—ฌ ์ฑ„ํŒ… ํ…œํ”Œ๋ฆฟ์— ์ „๋‹ฌํ•ฉ๋‹ˆ๋‹ค. LLM์€ ํ•จ์ˆ˜ ๋‚ด๋ถ€์˜ ์ฝ”๋“œ๋ฅผ ๋ณด์ง€ ๋ชปํ•ฉ๋‹ˆ๋‹ค. ๋‹ค์‹œ ๋งํ•ด, LLM์€ ํ•จ์ˆ˜๊ฐ€ ๊ธฐ์ˆ ์ ์œผ๋กœ ์–ด๋–ป๊ฒŒ ์ž‘๋™ํ•˜๋Š”์ง€๋Š” ์‹ ๊ฒฝ ์“ฐ์ง€ ์•Š๊ณ , ํ•จ์ˆ˜์˜ ์ •์˜์™€ ์ธ์ˆ˜๋งŒ ์ฐธ์กฐํ•ฉ๋‹ˆ๋‹ค.

ํ•จ์ˆ˜๊ฐ€ ์•ž์„œ ๋‚˜์—ด๋œ ๊ทœ์น™์„ ๋”ฐ๋ฅด๋ฉด, ๋‚ด๋ถ€์—์„œ JSON ์Šคํ‚ค๋งˆ๊ฐ€ ์ž๋™์œผ๋กœ ์ƒ์„ฑ๋ฉ๋‹ˆ๋‹ค. ํ•˜์ง€๋งŒ ๋” ๋‚˜์€ ๊ฐ€๋…์„ฑ์ด๋‚˜ ๋””๋ฒ„๊น…์„ ์œ„ํ•ด get_json_schema๋ฅผ ์‚ฌ์šฉํ•˜์—ฌ ์Šคํ‚ค๋งˆ๋ฅผ ์ˆ˜๋™์œผ๋กœ ๋ณ€ํ™˜ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

from transformers.utils import get_json_schema

def multiply(a: float, b: float):
    """
    ๋‘ ์ˆซ์ž๋ฅผ ๊ณฑํ•˜๋Š” ํ•จ์ˆ˜
    
    Args:
        a: ๊ณฑํ•  ์ฒซ ๋ฒˆ์งธ ์ˆซ์ž
        b: ๊ณฑํ•  ๋‘ ๋ฒˆ์งธ ์ˆซ์ž
    """
    return a * b

schema = get_json_schema(multiply)
print(schema)
{
  "type": "function", 
  "function": {
    "name": "multiply", 
    "description": "A function that multiplies two numbers", 
    "parameters": {
      "type": "object", 
      "properties": {
        "a": {
          "type": "number", 
          "description": "The first number to multiply"
        }, 
        "b": {
          "type": "number",
          "description": "The second number to multiply"
        }
      }, 
      "required": ["a", "b"]
    }
  }
}

์Šคํ‚ค๋งˆ๋ฅผ ํŽธ์ง‘ํ•˜๊ฑฐ๋‚˜ ์ฒ˜์Œ๋ถ€ํ„ฐ ์ง์ ‘ ์ž‘์„ฑํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ์ด๋ฅผ ํ†ตํ•ด ๋” ๋ณต์žกํ•œ ํ•จ์ˆ˜์— ๋Œ€ํ•œ ์ •ํ™•ํ•œ ์Šคํ‚ค๋งˆ๋ฅผ ์œ ์—ฐํ•˜๊ฒŒ ์ •์˜ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Warning

ํ•จ์ˆ˜ ์‹œ๊ทธ๋‹ˆ์ฒ˜๋ฅผ ๋‹จ์ˆœํ•˜๊ฒŒ ์œ ์ง€ํ•˜๊ณ  ์ธ์ˆ˜๋ฅผ ์ตœ์†Œํ•œ์œผ๋กœ ์œ ์ง€ํ•˜์„ธ์š”. ์ด๋Ÿฌํ•œ ํ•จ์ˆ˜๋Š” ์ค‘์ฒฉ๋œ ์ธ์ˆ˜๋ฅผ ๊ฐ€์ง„ ๋ณต์žกํ•œ ํ•จ์ˆ˜์— ๋น„ํ•ด ๋ชจ๋ธ์ด ๋” ์‰ฝ๊ฒŒ ์ดํ•ดํ•˜๊ณ  ์‚ฌ์šฉํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

์•„๋ž˜ ์˜ˆ์‹œ๋Š” ์Šคํ‚ค๋งˆ๋ฅผ ์ˆ˜๋™์œผ๋กœ ์ž‘์„ฑํ•œ ๋‹ค์Œ [~PreTrainedTokenizerBase.apply_chat_template]์— ์ „๋‹ฌํ•˜๋Š” ๋ฐฉ๋ฒ•์„ ๋ณด์—ฌ์ค๋‹ˆ๋‹ค.

# ์ธ์ˆ˜๋ฅผ ๋ฐ›์ง€ ์•Š๋Š” ๊ฐ„๋‹จํ•œ ํ•จ์ˆ˜
current_time = {
  "type": "function", 
  "function": {
    "name": "current_time",
    "description": "Get the current local time as a string.",
    "parameters": {
      'type': 'object',
      'properties': {}
    }
  }
}

# ๋‘ ๊ฐœ์˜ ์ˆซ์ž ์ธ์ˆ˜๋ฅผ ๋ฐ›๋Š” ๋” ์™„์ „ํ•œ ํ•จ์ˆ˜
multiply = {
  'type': 'function',
  'function': {
    'name': 'multiply',
    'description': 'A function that multiplies two numbers', 
    'parameters': {
      'type': 'object', 
      'properties': {
        'a': {
          'type': 'number',
          'description': 'The first number to multiply'
        }, 
        'b': {
          'type': 'number', 'description': 'The second number to multiply'
        }
      }, 
      'required': ['a', 'b']
    }
  }
}

model_input = tokenizer.apply_chat_template(
    messages,
    tools = [current_time, multiply]
)

RAGRAG

๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(Retrieval-augmented generation, RAG) ๋ชจ๋ธ์€ ์ฟผ๋ฆฌ๋ฅผ ๋ฐ˜ํ™˜ํ•˜๊ธฐ ์ „์— ๋ฌธ์„œ๋ฅผ ๊ฒ€์ƒ‰ํ•ด ์ถ”๊ฐ€ ์ •๋ณด๋ฅผ ์–ป์–ด ๋ชจ๋ธ์ด ๊ธฐ์กด์— ๊ฐ€์ง€๊ณ  ์žˆ๋˜ ์ง€์‹์„ ํ™•์žฅ์‹œํ‚ต๋‹ˆ๋‹ค. RAG ๋ชจ๋ธ์˜ ๊ฒฝ์šฐ, [~PreTrainedTokenizerBase.apply_chat_template]์— documents ๋งค๊ฐœ๋ณ€์ˆ˜๋ฅผ ์ถ”๊ฐ€ํ•˜์„ธ์š”. ์ด documents ๋งค๊ฐœ๋ณ€์ˆ˜๋Š” ๋ฌธ์„œ ๋ชฉ๋ก์ด์–ด์•ผ ํ•˜๋ฉฐ, ๊ฐ ๋ฌธ์„œ๋Š” title๊ณผ content ํ‚ค๋ฅผ ๊ฐ€์ง„ ๋‹จ์ผ ๋”•์…”๋„ˆ๋ฆฌ์—ฌ์•ผ ํ•ฉ๋‹ˆ๋‹ค.

Tip

RAG๋ฅผ ์œ„ํ•œ documents ๋งค๊ฐœ๋ณ€์ˆ˜๋Š” ํญ๋„“๊ฒŒ ์ง€์›๋˜์ง€ ์•Š์œผ๋ฉฐ ๋งŽ์€ ๋ชจ๋ธ๋“ค์ด documents๋ฅผ ๋ฌด์‹œํ•˜๋Š” ์ฑ„ํŒ… ํ…œํ”Œ๋ฆฟ์„ ๊ฐ€์ง€๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค. ๋ชจ๋ธ์ด documents๋ฅผ ์ง€์›ํ•˜๋Š”์ง€ ํ™•์ธํ•˜๋ ค๋ฉด ๋ชจ๋ธ ์นด๋“œ๋ฅผ ์ฝ๊ฑฐ๋‚˜ print(tokenizer.chat_template)๋ฅผ ์‹คํ–‰ํ•˜์—ฌ documents ํ‚ค๊ฐ€ ์žˆ๋Š”์ง€ ํ™•์ธํ•˜์„ธ์š”. Command-R๊ณผ Command-R+๋Š” ๋ชจ๋‘ RAG ์ฑ„ํŒ… ํ…œํ”Œ๋ฆฟ์—์„œ documents๋ฅผ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค.

๋ชจ๋ธ์— ์ „๋‹ฌํ•  ๋ฌธ์„œ ๋ชฉ๋ก์„ ์ƒ์„ฑํ•˜์„ธ์š”.

documents = [
    {
        "title": "The Moon: Our Age-Old Foe", 
        "text": "Man has always dreamed of destroying the moon. In this essay, I shall..."
    },
    {
        "title": "The Sun: Our Age-Old Friend",
        "text": "Although often underappreciated, the sun provides several notable benefits..."
    }
]

[~PreTrainedTokenizerBase.apply_chat_template]์—์„œ chat_template="rag"๋ฅผ ์„ค์ •ํ•˜๊ณ  ์‘๋‹ต์„ ์ƒ์„ฑํ•˜์„ธ์š”.

from transformers import AutoTokenizer, AutoModelForCausalLM

# ๋ชจ๋ธ๊ณผ ํ† ํฌ๋‚˜์ด์ € ๋กœ๋“œ
tokenizer = AutoTokenizer.from_pretrained("CohereForAI/c4ai-command-r-v01-4bit")
model = AutoModelForCausalLM.from_pretrained("CohereForAI/c4ai-command-r-v01-4bit", device_map="auto")
device = model.device # ๋ชจ๋ธ์„ ๊ฐ€์ ธ์˜จ ์žฅ์น˜ ํ™•์ธ

# ๋Œ€ํ™” ์ž…๋ ฅ ์ •์˜
conversation = [
    {"role": "user", "content": "What has Man always dreamed of?"}
]

input_ids = tokenizer.apply_chat_template(
    conversation=conversation,
    documents=documents,
    chat_template="rag",
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt").to(device)

# ์‘๋‹ต ์ƒ์„ฑ
generated_tokens = model.generate(
    input_ids,
    max_new_tokens=100,
    do_sample=True,
    temperature=0.3,
    )

# ์ƒ์„ฑ๋œ ํ…์ŠคํŠธ๋ฅผ ๋””์ฝ”๋”ฉํ•˜๊ณ  ์ƒ์„ฑ ํ”„๋กฌํ”„ํŠธ์™€ ํ•จ๊ป˜ ์ถœ๋ ฅ
generated_text = tokenizer.decode(generated_tokens[0])
print(generated_text)