huggingface--transformers
e06fe8e8c6
Secret Leaks / trufflehog (push) Failing after 1s
Build documentation / build (push) Failing after 1s
Build documentation / build_other_lang (push) Failing after 0s
CodeQL Security Analysis / CodeQL Analysis (push) Failing after 0s
PR CI / pr-ci (push) Failing after 1s
Slow tests on important models (on Push - A10) / Get all modified files (push) Failing after 1s
Slow tests on important models (on Push - A10) / Model CI (push) Has been skipped
Self-hosted runner (benchmark) / Benchmark (aws-g5-4xlarge-cache) (push) Has been cancelled
New model PR merged notification / Notify new model (push) Has been cancelled
Update Transformers metadata / build_and_package (push) Has been cancelled
101 ่ก
4.4 KiB
Markdown
101 ่ก
4.4 KiB
Markdown
<!--Copyright 2024 The HuggingFace Team. All rights reserved.
|
|
|
|
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with
|
|
the License. You may obtain a copy of the License at
|
|
|
|
http://www.apache.org/licenses/LICENSE-2.0
|
|
|
|
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on
|
|
an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the
|
|
specific language governing permissions and limitations under the License.
|
|
|
|
โ ๏ธ Note that this file is in Markdown but contain specific syntax for our doc-builder (similar to MDX) that may not be
|
|
rendered properly in your Markdown viewer.
|
|
|
|
-->
|
|
|
|
# GGUF์ Transformers์ ์ํธ์์ฉ [[gguf-and-interaction-with-transformers]]
|
|
|
|
GGUF ํ์ผ ํ์์ [GGML](https://github.com/ggerganov/ggml)๊ณผ ๊ทธ์ ์์กดํ๋ ๋ค๋ฅธ ๋ผ์ด๋ธ๋ฌ๋ฆฌ, ์๋ฅผ ๋ค์ด ๋งค์ฐ ์ธ๊ธฐ ์๋ [llama.cpp](https://github.com/ggerganov/llama.cpp)์ด๋ [whisper.cpp](https://github.com/ggerganov/whisper.cpp)์์ ์ถ๋ก ์ ์ํ ๋ชจ๋ธ์ ์ ์ฅํ๋๋ฐ ์ฌ์ฉ๋ฉ๋๋ค.
|
|
|
|
์ด ํ์ผ ํ์์ [Hugging Face Hub](https://huggingface.co/docs/hub/en/gguf)์์ ์ง์๋๋ฉฐ, ํ์ผ ๋ด์ ํ
์์ ๋ฉํ๋ฐ์ดํฐ๋ฅผ ์ ์ํ๊ฒ ๊ฒ์ฌํ ์ ์๋ ๊ธฐ๋ฅ์ ์ ๊ณตํฉ๋๋ค.
|
|
|
|
์ด ํ์์ "๋จ์ผ ํ์ผ ํ์(single-file-format)"์ผ๋ก ์ค๊ณ๋์์ผ๋ฉฐ, ํ๋์ ํ์ผ์ ์ค์ ์์ฑ, ํ ํฌ๋์ด์ ์ดํ, ๊ธฐํ ์์ฑ๋ฟ๋ง ์๋๋ผ ๋ชจ๋ธ์์ ๋ก๋๋๋ ๋ชจ๋ ํ
์๊ฐ ํฌํจ๋ฉ๋๋ค. ์ด ํ์ผ๋ค์ ํ์ผ์ ์์ํ ์ ํ์ ๋ฐ๋ผ ๋ค๋ฅธ ํ์์ผ๋ก ์ ๊ณต๋ฉ๋๋ค. ๋ค์ํ ์์ํ ์ ํ์ ๋ํ ๊ฐ๋ตํ ์ค๋ช
์ [์ฌ๊ธฐ](https://huggingface.co/docs/hub/en/gguf#quantization-types)์์ ํ์ธํ ์ ์์ต๋๋ค.
|
|
|
|
## Transformers ๋ด ์ง์ [[support-within-transformers]]
|
|
|
|
`transformers` ๋ด์์ `gguf` ํ์ผ์ ๋ก๋ํ ์ ์๋ ๊ธฐ๋ฅ์ ์ถ๊ฐํ์ฌ GGUF ๋ชจ๋ธ์ ์ถ๊ฐ ํ์ต/๋ฏธ์ธ ์กฐ์ ์ ์ ๊ณตํ ํ `ggml` ์ํ๊ณ์์ ๋ค์ ์ฌ์ฉํ ์ ์๋๋ก `gguf` ํ์ผ๋ก ๋ณํํ๋ ๊ธฐ๋ฅ์ ์ ๊ณตํฉ๋๋ค. ๋ชจ๋ธ์ ๋ก๋ํ ๋ ๋จผ์ FP32๋ก ์ญ์์ํํ ํ, PyTorch์์ ์ฌ์ฉํ ์ ์๋๋ก ๊ฐ์ค์น๋ฅผ ๋ก๋ํฉ๋๋ค.
|
|
|
|
> [!NOTE]
|
|
> ์ง์์ ์์ง ์ด๊ธฐ ๋จ๊ณ์ ์์ผ๋ฉฐ, ๋ค์ํ ์์ํ ์ ํ๊ณผ ๋ชจ๋ธ ์ํคํ
์ฒ์ ๋ํด ์ด๋ฅผ ๊ฐํํ๊ธฐ ์ํ ๊ธฐ์ฌ๋ฅผ ํ์ํฉ๋๋ค.
|
|
|
|
ํ์ฌ ์ง์๋๋ ๋ชจ๋ธ ์ํคํ
์ฒ์ ์์ํ ์ ํ์ ๋ค์๊ณผ ๊ฐ์ต๋๋ค:
|
|
|
|
### ์ง์๋๋ ์์ํ ์ ํ [[supported-quantization-types]]
|
|
|
|
์ด๊ธฐ์ ์ง์๋๋ ์์ํ ์ ํ์ Hub์์ ๊ณต์ ๋ ์ธ๊ธฐ ์๋ ์์ํ ํ์ผ์ ๋ฐ๋ผ ๊ฒฐ์ ๋์์ต๋๋ค.
|
|
|
|
- F32
|
|
- F16
|
|
- BF16
|
|
- Q4_0
|
|
- Q4_1
|
|
- Q5_0
|
|
- Q5_1
|
|
- Q8_0
|
|
- Q2_K
|
|
- Q3_K
|
|
- Q4_K
|
|
- Q5_K
|
|
- Q6_K
|
|
- IQ1_S
|
|
- IQ1_M
|
|
- IQ2_XXS
|
|
- IQ2_XS
|
|
- IQ2_S
|
|
- IQ3_XXS
|
|
- IQ3_S
|
|
- IQ4_XS
|
|
- IQ4_NL
|
|
|
|
> [!NOTE]
|
|
> GGUF ์ญ์์ํ๋ฅผ ์ง์ํ๋ ค๋ฉด `gguf>=0.10.0` ์ค์น๊ฐ ํ์ํฉ๋๋ค.
|
|
|
|
### ์ง์๋๋ ๋ชจ๋ธ ์ํคํ
์ฒ [[supported-model-architectures]]
|
|
|
|
ํ์ฌ ์ง์๋๋ ๋ชจ๋ธ ์ํคํ
์ฒ๋ Hub์์ ๋งค์ฐ ์ธ๊ธฐ๊ฐ ๋ง์ ์ํคํ
์ฒ๋ค๋ก ์ ํ๋์ด ์์ต๋๋ค:
|
|
|
|
- LLaMa
|
|
- Mistral
|
|
- Qwen2
|
|
- Qwen2Moe
|
|
- Phi3
|
|
- Bloom
|
|
|
|
## ์ฌ์ฉ ์์ [[example-usage]]
|
|
|
|
`transformers`์์ `gguf` ํ์ผ์ ๋ก๋ํ๋ ค๋ฉด `from_pretrained` ๋ฉ์๋์ `gguf_file` ์ธ์๋ฅผ ์ง์ ํด์ผ ํฉ๋๋ค. ๋์ผํ ํ์ผ์์ ํ ํฌ๋์ด์ ์ ๋ชจ๋ธ์ ๋ก๋ํ๋ ๋ฐฉ๋ฒ์ ๋ค์๊ณผ ๊ฐ์ต๋๋ค:
|
|
|
|
```python
|
|
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|
|
|
model_id = "TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF"
|
|
filename = "tinyllama-1.1b-chat-v1.0.Q6_K.gguf"
|
|
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
|
|
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
|
|
```
|
|
|
|
์ด์ PyTorch ์ํ๊ณ์์ ๋ชจ๋ธ์ ์์ํ๋์ง ์์ ์ ์ฒด ๋ฒ์ ์ ์ ๊ทผํ ์ ์์ผ๋ฉฐ, ๋ค๋ฅธ ์ฌ๋ฌ ๋๊ตฌ๋ค๊ณผ ๊ฒฐํฉํ์ฌ ์ฌ์ฉํ ์ ์์ต๋๋ค.
|
|
|
|
`gguf` ํ์ผ๋ก ๋ค์ ๋ณํํ๋ ค๋ฉด llama.cpp์ [`convert-hf-to-gguf.py`](https://github.com/ggerganov/llama.cpp/blob/master/convert_hf_to_gguf.py)๋ฅผ ์ฌ์ฉํ๋ ๊ฒ์ ๊ถ์ฅํฉ๋๋ค.
|
|
|
|
์์ ์คํฌ๋ฆฝํธ๋ฅผ ์๋ฃํ์ฌ ๋ชจ๋ธ์ ์ ์ฅํ๊ณ ๋ค์ `gguf`๋ก ๋ด๋ณด๋ด๋ ๋ฐฉ๋ฒ์ ๋ค์๊ณผ ๊ฐ์ต๋๋ค:
|
|
|
|
```python
|
|
tokenizer.save_pretrained('directory')
|
|
model.save_pretrained('directory')
|
|
|
|
!python ${path_to_llama_cpp}/convert-hf-to-gguf.py ${directory}
|
|
```
|