> [!NOTE] > 本文档由 WeHub 基于上游 README 翻译整理,属于社区翻译,非官方中文文档。 > [English](./README.en.md) · [原始项目](https://github.com/DrewThomasson/ebook2audiobook) · [上游 README](https://github.com/DrewThomasson/ebook2audiobook/blob/HEAD/README.md) > 原作者、版权与许可证归属以原始项目及本仓库 LICENSE 文件为准。 # 📚 ebook2audiobook (E2A) CPU/GPU 电子书转有声书转换器,支持章节与元数据
采用先进的 TTS 引擎及更多功能。
支持声音克隆和 1158 种语言! > [!IMPORTANT] **本工具仅适用于无 DRM、合法获取的电子书。**
作者不对本软件的任何滥用行为或由此产生的法律后果负责。
请负责任地使用本工具,并遵守所有适用法律。 [![Discord](https://dcbadge.limes.pink/api/server/https://discord.gg/63Tv3F65k6)](https://discord.gg/63Tv3F65k6) ### 感谢支持 ebook2audiobook 开发者! [![Ko-Fi](https://img.shields.io/badge/Ko--fi-F16061?style=for-the-badge&logo=ko-fi&logoColor=white)](https://ko-fi.com/athomasson2) ### 本地运行 [![Quick Start](https://img.shields.io/badge/Quick%20Start-blue?style=for-the-badge)](#instructions) [![Docker Build](https://github.com/DrewThomasson/ebook2audiobook/actions/workflows/Docker-Build.yml/badge.svg)](https://github.com/DrewThomasson/ebook2audiobook/actions/workflows/Docker-Build.yml) [![Download](https://img.shields.io/badge/Download-Now-blue.svg)](https://github.com/DrewThomasson/ebook2audiobook/releases/latest) Platform Docker Pull Count ### 远程运行 [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-Spaces-yellow?style=flat&logo=huggingface)](https://huggingface.co/spaces/drewThomasson/ebook2audiobook) [![Free Google Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/DrewThomasson/ebook2audiobook/blob/main/Notebooks/colab_ebook2audiobook.ipynb) [![Kaggle](https://img.shields.io/badge/Kaggle-035a7d?style=flat&logo=kaggle&logoColor=white)](https://github.com/Rihcus/ebook2audiobookXTTS/blob/main/Notebooks/kaggle-ebook2audiobook.ipynb) #### GUI 界面 ![demo_web_gui](assets/demo_web_gui.gif)
点击查看 Web GUI 截图 GUI Screen 1 GUI Screen 2 GUI Screen 3
## 演示 **新默认语音演示** https://github.com/user-attachments/assets/750035dc-e355-46f1-9286-05c1d9e88cea
更多演示 **ASMR 语音** https://github.com/user-attachments/assets/68eee9a1-6f71-4903-aacd-47397e47e422 **雨天语音** https://github.com/user-attachments/assets/d25034d9-c77f-43a9-8f14-0d167172b080 **Scarlett 语音** https://github.com/user-attachments/assets/b12009ee-ec0d-45ce-a1ef-b3a52b9f8693 **David Attenborough 语音** https://github.com/user-attachments/assets/81c4baad-117e-4db5-ac86-efc2b7fea921 **示例** ![Example](https://github.com/DrewThomasson/VoxNovel/blob/dc5197dff97252fa44c391dc0596902d71278a88/readme_files/example_in_app.jpeg)
## README.md ## 目录 - [ebook2audiobook](#-ebook2audiobook) - [功能](#features) - [GUI 界面](#gui-interface) - [演示](#demos) - [支持的语言](#supported-languages) - [最低要求](#hardware-requirements) - [用法](#instructions) - [本地运行](#instructions) - [启动 Gradio Web 界面](#instructions) - [基本无头模式用法](#basic-usage) - [无头模式自定义 XTTS 模型用法](#example-of-custom-model-zip-upload) - [帮助命令输出](#help-command-output) - [远程运行](#run-remotely) - [Docker](#docker) - [运行步骤](#docker) - [克隆语音](#cloned-voices) - [微调 TTS 模型](#fine-tuned-tts-models) - [微调 TTS 模型合集](#fine-tuned-tts-collection) - [训练 XTTSv2](#fine-tune-your-own-xttsv2-model) - [支持的电子书格式](#supported-ebook-formats) - [输出格式](#output-and-process-formats) - [回退到旧版本](#reverting-to-older-versions) - [常见问题](#common-issues) - [特别感谢](#special-thanks) - [目录](#table-of-contents) ## 功能 - 🔧 **支持的 TTS 引擎**:`XTTSv2`, `Bark`, `Fairseq`, `VITS`, `Tacotron2`, `Tortoise`, `GlowTTS`, `YourTTS` - 📚 **转换多种文件格式**:`.epub`, `.mobi`, `.azw3`, `.fb2`, `.lrf`, `.rb`, `.snb`, `.tcr`, `.pdf`, `.txt`, `.rtf`, `.doc`, `.docx`, `.html`, `.odt`, `.azw`, `.tiff`, `.tif`, `.png`, `.jpg`, `.jpeg`, `.bmp`, `.zip` - 💻 **TextArea** 可直接将短文本转换为音频 - 🔍 对以图像形式呈现文本页面的文件进行 **OCR 扫描** - 🔊 **高质量文本转语音**,从近实时到接近真人语音 - 🗣️ **可选声音克隆**,使用你自己的语音文件 - 🌐 **支持 1158 种语言**([支持语言列表](https://dl.fbaipublicfiles.com/mms/tts/all-tts-languages.html)) - 💻 **低资源友好** — 最低可在 **2 GB RAM / 1 GB VRAM** 上运行 - 🎵 **有声书输出格式**:单声道或立体声 `aac`, `flac`, `mp3`, `m4b`, `m4a`, `mp4`, `mov`, `ogg`, `wav`, `webm` - 🧠 **支持 SML 标签** — 精细控制停顿、暂停、语音切换等([见下文](#sml-tags-available)) - 🧩 **可选自定义模型**,使用你自己训练的模型(XTTSv2、VITS、FAIRSEQ、PIPER,其他可按需支持) - 🎛️ **微调预设模型**,由 E2A 团队训练
(如需更多微调模型,或希望将你的模型分享到官方预设列表,请联系我们) ## 硬件要求 - 最低 2GB RAM,建议 8GB。 - 最低 1GB VRAM,建议 4GB。 - 在 Windows 上运行时需启用虚拟化(仅 Docker)。 - CPU、XPU(Intel、AMD、ARM)*。 - CUDA、ROCm、JETSON - MPS(Apple Silicon CPU) *现代 TTS 引擎在 CPU 上非常慢,因此请使用较低质量的 TTS,如 YourTTS、Tacotron2 等。 ## 支持的语言 | **Arabic (ar)** | **Chinese (zh)** | **English (en)** | **Spanish (es)** | |:------------------:|:------------------:|:------------------:|:------------------:| | **French (fr)** | **German (de)** | **Italian (it)** | **Portuguese (pt)** | | **Polish (pl)** | **Turkish (tr)** | **Russian (ru)** | **Dutch (nl)** | | **Czech (cs)** | **Japanese (ja)** | **Hindi (hi)** | **Bengali (bn)** | | **Hungarian (hu)** | **Korean (ko)** | **Vietnamese (vi)**| **Swedish (sv)** | | **Persian (fa)** | **Yoruba (yo)** | **Swahili (sw)** | **Indonesian (id)**| | **Slovak (sk)** | **Croatian (hr)** | **Tamil (ta)** | **Danish (da)** | - [**此处另有 +1130 种语言和方言**](https://dl.fbaipublicfiles.com/mms/tts/all-tts-languages.html) ## 支持的电子书格式 - `.epub`, `.pdf`, `.mobi`, `.txt`, `.html`, `.rtf`, `.chm`, `.lit`, `.pdb`, `.fb2`, `.odt`, `.cbr`, `.cbz`, `.prc`, `.lrf`, `.pml`, `.snb`, `.cbc`, `.rb`, `.tcr` - **最佳效果**:使用 `.epub` 或 `.mobi` 可自动检测章节 ## 输出与处理格式 - `.m4b`, `.m4a`, `.mp4`, `.webm`, `.mov`, `.mp3`, `.flac`, `.wav`, `.ogg`, `.aac` - 处理格式可在 lib/conf.py 中修改 ## 可用的 SML 标签 - `[break]` — 静音(随机范围 **0.3–0.6 秒**) - `[pause]` — 静音(随机范围 **1.0–1.6 秒**) - `[pause:N]` — 固定暂停(**N 秒**) - `[voice:/path/to/voice/file]...[/voice]` — 从默认或 GUI/CLI 所选语音切换语音 **查看我们专门用于自动向电子书添加 SML 的其他仓库 -> [E2A-SML](./components/E2A-SML)** > [!IMPORTANT] **在提交安装或 bug 问题之前,请仔细搜索已打开和已关闭的 issues 标签页
以确保你的问题尚未存在。** >[!NOTE] **EPUB 格式缺乏任何标准结构,例如章节、段落、前言等。
因此你应首先手动删除任何不想转换为音频的文本。** ### 使用说明 1. **克隆仓库** ```bash git clone https://github.com/DrewThomasson/ebook2audiobook.git cd ebook2audiobook ``` 2. **安装 / 运行 ebook2audiobook**: - **Linux/MacOS** ```bash ./ebook2audiobook.command ``` MacOS 用户请注意:将安装 homebrew 以安装缺失的程序。 - **Mac 启动器** 双击 `Mac Ebook2Audiobook Launcher.command` - **Windows** ```bash ebook2audiobook.cmd ``` 或 双击 `ebook2audiobook.cmd` Windows 用户请注意:将安装 scoop,以便在没有管理员权限的情况下安装缺失的程序。 1. **打开 Web 应用**:点击终端中提供的 URL 以访问 Web 应用并转换 eBook。`http://localhost:7860/` 2. **公开链接**: `./ebook2audiobook.command --share` (Linux/MacOS) `ebook2audiobook.cmd --share` (Windows) `python app.py --share` (all OS) > [!IMPORTANT] **如果脚本停止后再次运行,你需要刷新 gradio GUI 界面
以便网页重新连接到新的连接 socket。** ### 基本用法 - **Linux/MacOS**: ```bash ./ebook2audiobook.command --headless --ebook --voice --language ``` - **Windows** ```bash ebook2audiobook.cmd --headless --ebook --voice --language ``` - **[--ebook]**:你的 eBook 文件路径 - **[--voice]**:语音克隆文件路径(可选) - **[--language]**:ISO-639-3 语言代码(例如:ita 表示意大利语,eng 表示英语,deu 表示德语……)。
默认语言为 eng,对于在 ./lib/lang.py 中设置的默认语言,--language 为可选参数。
也支持 ISO-639-1 双字母代码。 ### 自定义模型 Zip 上传示例 (必须是包含必需模型文件的 .zip 文件。以 XTTSv2 为例:config.json、model.pth、vocab.json 和 ref.wav) - **Linux/MacOS** ```bash ./ebook2audiobook.command --headless --ebook --language --custom_model ``` - **Windows** ```bash ebook2audiobook.cmd --headless --ebook --language --custom_model ``` 注意:自定义模型的 ref.wav 始终是用于转换时所选的语音 - ****:`model_name.zip` 文件的路径, 该文件必须(根据 TTS 引擎)包含所有必需文件
(参见 ./lib/models.py)。 ### 详细指南及全部可用参数列表 - **Linux/MacOS** ```bash ./ebook2audiobook.command --help ``` - **Windows** ```bash ebook2audiobook.cmd --help ``` - **或适用于所有操作系统** ```python app.py --help ``` ```bash usage: app.py [-h] [--session SESSION] [--share] [--headless] [--ebook EBOOK] [--ebooks_dir EBOOKS_DIR] [--language LANGUAGE] [--voice VOICE] [--voice_map VOICE_MAP] [--device {CPU,CUDA,MPS,ROCM,XPU,JETSON}] [--tts_engine {XTTS,BARK,VITS,FAIRSEQ,TACOTRON,YOURTTS,xtts,bark,vits,fairseq,tacotron,yourtts}] [--custom_model CUSTOM_MODEL] [--fine_tuned FINE_TUNED] [--output_format OUTPUT_FORMAT] [--output_channel OUTPUT_CHANNEL] [--temperature TEMPERATURE] [--length_penalty LENGTH_PENALTY] [--num_beams NUM_BEAMS] [--repetition_penalty REPETITION_PENALTY] [--top_k TOP_K] [--top_p TOP_P] [--speed SPEED] [--enable_text_splitting] [--text_temp TEXT_TEMP] [--waveform_temp WAVEFORM_TEMP] [--output_dir OUTPUT_DIR] [--version] Convert eBooks to Audiobooks using a Text-to-Speech model. You can either launch the Gradio interface or run the script in headless mode for direct conversion. options: -h, --help show this help message and exit --session SESSION Session to resume the conversion in case of interruption, crash, or reuse of custom models and custom cloning voices. **** The following option is for gradio/gui mode only: --share (Optional) Enable a public shareable Gradio link. **** The following options are for --headless mode only: --headless Run the script in headless mode --ebook EBOOK Path to the ebook file for conversion. Cannot be used when --ebooks_dir is present. --ebooks_dir EBOOKS_DIR Relative or absolute path of the directory containing the files to convert. Cannot be used when --ebook is present. --text TEXT Raw text for conversion. Cannot be used when --ebook or --ebooks_dir is present. --language LANGUAGE Language of the e-book. Default language is set in ./lib/lang.py sed as default if not present. All compatible language codes are in ./lib/lang.py optional parameters: --translate ISO3 (Optional) Translate ebook to a target language (ISO 639-3 code, e.g. eng, fra, deu) before TTS synthesis. Uses argostranslate. The target language becomes the effective TTS language for the run. A copy of the source ebook is made with the _ suffix so translated and non-translated outputs stay isolated (independent process folder, audio chunks, and final file). --voice VOICE (Optional) Path to the voice cloning file for TTS engine. Uses the default voice if not present. --voice_map VOICE_MAP (Optional, --ebooks_dir only) Path to a JSON file mapping ebook path -> voice path. Each entry overrides --voice for that specific ebook. Missing/null entries fall back to --voice. Keys may be absolute paths or basenames. Example: {"book1.epub": "/voices/eng/adult/female/alice.wav", "/abs/path/book2.epub": null} --device {CPU,CUDA,MPS,ROCM,XPU,JETSON} (Optional) Processor unit type for the conversion. Default is set in ./lib/conf.py if not present. Fall back to CPU if CUDA or MPS is not available. --tts_engine {XTTS,BARK,VITS,FAIRSEQ,TACOTRON,YOURTTS,xtts,bark,vits,fairseq,tacotron,yourtts} (Optional) Preferred TTS engine (available are: ['XTTS', 'BARK', 'VITS', 'FAIRSEQ', 'TACOTRON', 'YOURTTS', 'xtts', 'bark', 'vits', 'fairseq', 'tacotron', 'yourtts']. Default depends on the selected language. The tts engine should be compatible with the chosen language --custom_model CUSTOM_MODEL (Optional) Path to the custom model zip file cntaining mandatory model files. Please refer to ./lib/models.py --fine_tuned FINE_TUNED (Optional) Fine tuned model path. Default is builtin model. --output_format OUTPUT_FORMAT (Optional) Output audio format. Default is m4b set in ./lib/conf.py --output_channel OUTPUT_CHANNEL (Optional) Output audio channel. Default is mono set in ./lib/conf.py --temperature TEMPERATURE (xtts only, optional) Temperature for the model. Default to config.json model. Higher temperatures lead to more creative outputs. --length_penalty LENGTH_PENALTY (xtts only, optional) A length penalty applied to the autoregressive decoder. Default to config.json model. Not applied to custom models. --num_beams NUM_BEAMS (xtts only, optional) Controls how many alternative sequences the model explores. Must be equal or greater than length penalty. Default to config.json model. --repetition_penalty REPETITION_PENALTY (xtts only, optional) A penalty that prevents the autoregressive decoder from repeating itself. Default to config.json model. --top_k TOP_K (xtts only, optional) Top-k sampling. Lower values mean more likely outputs and increased audio generation speed. Default to config.json model. --top_p TOP_P (xtts only, optional) Top-p sampling. Lower values mean more likely outputs and increased audio generation speed. Default to config.json model. --speed SPEED (xtts only, optional) Speed factor for the speech generation. Default to config.json model. --enable_text_splitting (xtts only, optional) Enable TTS text splitting. This option is known to not be very efficient. Default to config.json model. --text_temp TEXT_TEMP (bark only, optional) Text Temperature for the model. Default to config.json model. --waveform_temp WAVEFORM_TEMP (bark only, optional) Waveform Temperature for the model. Default to config.json model. --output_dir OUTPUT_DIR (Optional) Path to the output directory. Default is set in ./lib/conf.py --version Show the version of the script and exit Example usage: Windows: Gradio/GUI: ebook2audiobook.cmd Headless mode: ebook2audiobook.cmd --headless --ebook '/path/to/file' --language eng Linux/Mac: Gradio/GUI: ./ebook2audiobook.command Headless mode: ./ebook2audiobook.command --headless --ebook '/path/to/file' --language eng SML tags available: [break] — silence (random range **0.3–0.6 sec.**) [pause] — silence (random range **1.0–1.6 sec.**) [pause:N] — fixed pause (**N sec.**) [voice:/path/to/voice/file]...[/voice] — switch voice from default or selected voice from GUI/CLI ``` 注意:在 gradio/gui 模式下,要取消正在进行的转换,只需点击电子书上传组件上的 [X]。 提示:如需更长停顿,可添加 '[pause:3]' 表示 3 秒等。 ### Docker 1. **克隆仓库(Clone the Repository)**: ```bash git clone https://github.com/DrewThomasson/ebook2audiobook.git cd ebook2audiobook ``` 2. **构建容器** ```bash Windows: Docker: ebook2audiobook.cmd --script_mode build_docker Docker Compose: ebook2audiobook.cmd --script_mode build_docker --docker_mode compose Podman Compose: ebook2audiobook.cmd --script_mode build_docker --docker_mode podman Linux/Mac Docker: ./ebook2audiobook.command --script_mode build_docker Docker Compose ./ebook2audiobook.command --script_mode build_docker --docker_mode compose Podman Compose: ./ebook2audiobook.command --script_mode build_docker --docker_mode podman ``` 4. **运行容器:** ```bash Docker run image: Gradio/GUI: CPU: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" --rm -it -p 7860:7860 athomasson2/ebook2audiobook:cpu CUDA: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" --gpus all --rm -it -p 7860:7860 athomasson2/ebook2audiobook:cu[118/122/124/126 etc..] ROCM: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" --device=/dev/kfd --device=/dev/dri --rm -it -p 7860:7860 athomasson2/ebook2audiobook:rocm[6.0/6.1/6.4 etc..] XPU: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" --device=/dev/dri --rm -it -p 7860:7860 athomasson2/ebook2audiobook:xpu JETSON: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" --runtime nvidia --rm -it -p 7860:7860 athomasson2/ebook2audiobook:jetson[51/60/61 etc...] Headless mode: CPU: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" -v "/my/real/ebooks/folder/absolute/path:/app/another_ebook_folder" --rm -it -p 7860:7860 ebook2audiobook:cpu --headless --ebook "/app/another_ebook_folder/myfile.pdf" [--voice /app/my/voicepath/voice.mp3 etc..] CUDA: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" -v "/my/real/ebooks/folder/absolute/path:/app/another_ebook_folder" --gpus all --rm -it -p 7860:7860 ebook2audiobook:cu[118/122/124/126 etc..] --headless --ebook "/app/another_ebook_folder/myfile.pdf" [--voice /app/my/voicepath/voice.mp3 etc..] ROCM: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" -v "/my/real/ebooks/folder/absolute/path:/app/another_ebook_folder" --device=/dev/kfd --device=/dev/dri --rm -it -p 7860:7860 ebook2audiobook:rocm[6.0/6.1/6.4 etc.] --headless --ebook "/app/another_ebook_folder/myfile.pdf" [--voice /app/my/voicepath/voice.mp3 etc..] XPU: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" -v "/my/real/ebooks/folder/absolute/path:/app/another_ebook_folder" --device=/dev/dri --rm -it -p 7860:7860 ebook2audiobook:xpu --headless --ebook "/app/another_ebook_folder/myfile.pdf" [--voice /app/my/voicepath/voice.mp3 etc..] JETSON: docker run -v "./ebooks:/app/ebooks" -v "./audiobooks:/app/audiobooks" -v "./models:/app/models" -v "./voices:/app/voices" -v "./tmp:/app/tmp" -v "/my/real/ebooks/folder/absolute/path:/app/another_ebook_folder" --runtime nvidia --rm -it -p 7860:7860 ebook2audiobook:jetson[51/60/61 etc.] --headless --ebook "/app/another_ebook_folder/myfile.pdf" [--voice /app/my/voicepath/voice.mp3 etc..] Docker Compose (i.e. cuda 12.8: Run Gradio GUI: DEVICE_TAG=cu128 docker compose --profile gpu up --no-log-prefix Run Headless mode: DEVICE_TAG=cu128 docker compose --profile gpu run --rm ebook2audiobook --headless --ebook "/app/ebooks/myfile.pdf" --voice /app/voices/eng/adult/female/some_voice.wav etc.. Podman Compose (i.e. cuda 12.8: Run Gradio GUI: DEVICE_TAG=cu128 podman-compose -f podman-compose.yml --profile gpu up Run Headless mode: DEVICE_TAG=cu128 podman-compose -f podman-compose.yml --profile gpu run --rm ebook2audiobook-gpu --headless --ebook "/app/ebooks/myfile.pdf" --voice /app/voices/eng/adult/female/some_voice.wav etc.. ``` - 注意:Docker 中未暴露 MPS,因此必须使用 CPU ## 克隆音色(Cloned Voices) 你可以上传任意支持音频格式的语音音频,理想时长约为 1 至 5 分钟。 录音背景嘈杂或有背景音乐也没关系 —— E2A 会为你清理人声。 内置克隆音色列表主要为英语。如需将其他语言的音色正式 添加到列表中,请联系我们,审核通过后会添加。 ## 微调 TTS 模型(Fine Tuned TTS models) #### 微调你自己的 XTTSv2 模型 [Universal_TTS_Finetune](./components/Universal_TTS_Finetune) [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-Spaces-yellow?style=flat&logo=huggingface)](https://huggingface.co/spaces/drewThomasson/xtts-finetune-webui-gpu) [![Kaggle](https://img.shields.io/badge/Kaggle-035a7d?style=flat&logo=kaggle&logoColor=white)](https://github.com/DrewThomasson/ebook2audiobook/blob/v25/Notebooks/finetune/xtts/kaggle-xtts-finetune-webui-gradio-gui.ipynb) [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/DrewThomasson/ebook2audiobook/blob/v25/Notebooks/finetune/xtts/colab_xtts_finetune_webui.ipynb) #### 训练数据降噪 [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-Spaces-yellow?style=flat&logo=huggingface)](https://huggingface.co/spaces/drewThomasson/DeepFilterNet2_no_limit) [![GitHub Repo](https://img.shields.io/badge/DeepFilterNet-181717?logo=github)](https://github.com/Rikorose/DeepFilterNet) ### 微调 TTS 合集 [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-Models-yellow?style=flat&logo=huggingface)](https://huggingface.co/drewThomasson/fineTunedTTSModels/tree/main) 对于 XTTSv2 自定义模型,必须提供参考音色的 ref 音频片段: ## 自定义你的 Ebook2Audiobook 你可以自由修改 libs/conf.py 以添加或移除所需设置。若打算这样做,请先 复制一份原始 conf.py,以便在每次 ebook2audiobook 更新时备份你修改过的 conf.py 并 换回原始文件。对 models.py 也应采用相同流程。若希望将自己的自定义模型 作为官方 ebook2audiobook 微调模型,请联系我们,我们会将其添加到预设列表。 ## 回退到旧版本 发布版本见 -> [此处](https://github.com/DrewThomasson/ebook2audiobook/releases) ```bash git checkout tags/VERSION_NUM # Locally/Compose -> Example: git checkout tags/v25.7.7 ``` ## 常见问题: - 我的 NVIDIA/ROCm/XPU/MPS GPU 未被检测到?? -> [GPU ISSUES Wiki Page](https://github.com/DrewThomasson/ebook2audiobook/wiki/GPU-ISSUES) - CPU 较慢(在多核服务器 CPU 上表现更好),而 GPU 可实现接近实时的转换。 [相关讨论](https://github.com/DrewThomasson/ebook2audiobook/discussions/19#discussioncomment-10879846) (不过它不支持零样本音色克隆(zero-shot voice cloning),且是 Siri 级音质,但在 CPU 上快得多)。 - 「我遇到依赖问题」—— 直接使用 Docker,它完全自包含且有无头模式, 在 docker run 命令末尾添加 `--help` 参数可获取更多信息。 - 「我遇到音频被截断的问题!」—— 请务必为此提交 ISSUE, 我们并非精通所有语言,需要用户的建议来微调分句逻辑。😊 ## ***** 路线图 ***** - 所有功能均开放公众贡献 ⭐ - 欢迎讲任何受支持语言的朋友帮助我们改进模型 ⭐ - [x] 在开始转换前预览区块/章节 - [ ] 按已转换的句子进行编辑,以实现精确文本修改 - [x] 集成 SML 标签,用于语音、停顿、break 及更多调整 - [x] 多语言的 -h -help 参数说明 - [x] 支持 PDF / JPG / BMP / PNG / TIFF 的 OCR 扫描 - [x] Notebooks 文件夹 [在此讨论](https://github.com/DrewThomasson/ebook2audiobookXTTS/issues/5#issuecomment-2408773254) - [x] 使中文文本切分不拆分词语,并改进停顿时间 [在此讨论](https://github.com/DrewThomasson/ebook2audiobookXTTS/issues/18#issuecomment-2401154894) - [x] Dockerfile - [x] Docker compose - [x] Podman compose - [x] Kaggle Notebook - [x] Google Colab Notebook - [ ] Audiobookshelf 集成 - [ ] [开发 iOS 应用](https://github.com/DrewThomasson/ebook2audiobook/pull/35#issuecomment-2496495212) - [ ] [开发 Android 应用](https://github.com/DrewThomasson/ebook2audiobook/pull/35#issuecomment-2496495212) #### 额外选项 - [x] 电子书翻译选项 - [x] 输出格式选择 - [x] 批量电子书文件夹 - [x] 多进程转换 - [x] 批量电子书文件夹转换 - [x] GPU 设备检测 - [x] 对任意参考音频进行降噪,用于上传语音克隆, - [x] 自定义模型上传(目前仅支持 XTTSv2,可按需扩展) - [ ] 为 xttsv2、fairseq、vits、piper 等至少添加欧洲葡萄牙语语言模型(欢迎协助) - [ ] 为 xttsv2、fairseq、vits、piper 等至少添加信德语语言模型(欢迎协助) #### TTS 引擎 - [x] XTTSv2 - [x] Bark - [x] Fairseq - [x] VITS - [x] Tacotron2 - [x] YourTTS - [x] Tortoise - [x] GlowTTS - [x] Piper - [ ] GPT-SoVITS (https://github.com/RVC-Boss/GPT-SoVITS) - [ ] OpenVoice (https://github.com/myshell-ai/OpenVoice) - [ ] fish-speech (https://github.com/fishaudio/fish-speech) - [ ] ChatTTS (https://github.com/2noise/ChatTTS) - [ ] CosyVoice (https://github.com/FunAudioLLM/CosyVoice) - [ ] F5-TTS (https://github.com/swivid/f5-tts) - [ ] chatterbox (https://github.com/resemble-ai/chatterbox) - [ ] Supertonic (https://github.com/supertone-inc/supertonic) - [ ] Spark-TTS (https://github.com/sparkaudio/spark-tts) - [ ] index-tts (https://github.com/index-tts/index-tts) - [ ] MeloTTS (https://github.com/myshell-ai/MeloTTS) - [ ] Kokoro-TTS (https://github.com/hexgrad/kokoro) - [ ] OmniVoice (https://github.com/k2-fsa/OmniVoice) - [ ] Zonos (https://github.com/Zyphra/Zonos) - [ ] Style-TTS2 (https://github.com/yl4579/StyleTTS2) - [ ] Orpheus-TTS (https://github.com/canopyai/Orpheus-TTS) - [ ] NewTTS (https://github.com/neuphonic/neutts?tab=readme-ov-file) - [ ] VIbeVoice (https://github.com/vibevoice-community/VibeVoice) - [ ] Qwen3-TTS (https://huggingface.co/spaces/Qwen/Qwen3-TTS) #### README 翻译 - [x] Arabic (ara) - [x] Chinese (zho) - [x] English (eng) - [x] Spanish (spa) - [x] French (fra) - [x] German (deu) - [x] Italian (ita) - [x] Portuguese (por) - [x] Polish (pol) - [x] Turkish (tur) - [x] Russian (rus) - [x] Dutch (nld) - [x] Czech (ces) - [x] Japanese (jpn) - [x] Hindi (hin) - [x] Bengali (ben) - [x] Hungarian (hun) - [x] Korean (kor) - [x] Vietnamese (vie) - [x] Swedish (swe) - [x] Persian (fas) - [x] Yoruba (yor) - [x] Swahili (swa) - [x] Indonesian (ind) - [x] Slovak (slk) - [x] Croatian (hrv) #### 🐍 操作系统兼容性 - [x] 🍎 Mac Intel x86 - [x] 🪟 Windows x86 - [x] 🐧 Linux x86 - [x] 🖥️🍏 Apple Silicon Mac - [x] 🪟💪 ARM Windows - [x] 🐧💪 ARM Linux ********** ## 用于训练模型等的额外进阶功能(一条简单命令即可支持所有 Coqui-tts 模型和 piper-tts) - 有关此功能的更多信息,请联系 @DrewThomasson,他目前正在开发此项工作,[进行中的仓库在此](https://github.com/DrewThomasson/Universal_TTS_Finetune) - [ ] 为所有 coqui-tts 模型制作易于使用的训练 GUI,采用 ljspeech 格式训练方案 [coqui tts 提供的方案在此](https://github.com/coqui-ai/TTS/tree/dev/recipes/ljspeech) ## 供贡献者参考的 Python 代码规范化说明 - 代码之间不留空行,函数和类之间除外。 - 除 dict() 和 json 外,所有键均使用单引号。dict['key'] 始终使用单引号调用 - 4 空格缩进,绝不使用 tab - 所有函数及其参数声明和返回值均需严格类型标注 - 参数与其类型标注之间不留空格,函数、“->” 与返回值之间不留空格 示例: ```python import json from typing import Optional def get_user(user_id:int, users:list[dict])->Optional[dict]: for user in users: if user['id'] == user_id: return user return None def summarize(user:dict)->str: return f"User {user['name']} is {'active' if user['is_active'] else 'inactive'}." def to_json(user:dict)->str: return json.dumps({"id": user['id'], "name": user['name'], "email": user['email']}) users:list = [ dict(id=1, name="alice", email="alice@example.com", role="admin", is_active=True), dict(id=2, name="bob", email="bob@example.com", role="editor", is_active=False), dict(id=3, name="carol", email="carol@example.com", role="viewer", is_active=True), ] config = { "max_users": 100, "default_role": "viewer", "allow_signup": True, } roles = ['admin', 'editor', 'viewer'] found = get_user(1, users) if found: print(summarize(found)) print(found['email']) print(to_json(found)) if config['default_role'] in roles: print(config['default_role']) ``` ## 征集用于 Beta 测试的硬件捐赠 我们接受各类硬件以测试我们的开发,例如: - 支持 CUDA >= 11.8 的 Nvidia 显卡 - Intel XPU 显卡 - 支持 ROCm >=5.7 的 AMD ROCm 显卡 @DrewThomasson 如果你想提供任何帮助!😃 ## 特别感谢 感谢所有资金和代码贡献者,每一份贡献与建议都有助于提升 E2A 的质量。