mirror of
https://github.com/xinnan-tech/xiaozhi-esp32-server.git
synced 2026-07-25 16:43:55 +08:00
update:更新说明文档 (#649)
* update:添加设备注册验证码接口 * update:更新说明文档 * update:优化说明文档 * update:更新说明文档
This commit is contained in:
+199
-195
@@ -1,8 +1,29 @@
|
||||
[](https://github.com/xinnan-tech/xiaozhi-esp32-server)
|
||||
|
||||
<center>
|
||||
<h1>Xiaozhi Backend Server xiaozhi-esp32-server</h1>
|
||||
</center>
|
||||
|
||||
[](https://github.com/xinnan-tech/xiaozhi-esp32-server)
|
||||
<p align="center">
|
||||
This project provides backend services for the open-source smart hardware project
|
||||
<a href="https://github.com/78/xiaozhi-esp32">xiaozhi-esp32</a><br/>
|
||||
Implemented in Python according to the <a href="https://ccnphfhqs21z.feishu.cn/wiki/M0XiwldO9iJwHikpXD5cEx71nKh">Xiaozhi Communication Protocol</a><br/>
|
||||
Helping you quickly set up your Xiaozhi server
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="./README.md">简体中文</a>
|
||||
· English
|
||||
· <a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/releases">Changelog</a>
|
||||
· <a href="./docs/Deployment.md">Deployment Guide</a>
|
||||
· <a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/issues">Report Issues</a>
|
||||
</p>
|
||||
<p align="center">
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/releases">
|
||||
<img alt="GitHub Contributors" src="https://img.shields.io/github/v/release/xinnan-tech/xiaozhi-esp32-server?logo=docker" />
|
||||
</a>
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/graphs/contributors">
|
||||
<img alt="GitHub Contributors" src="https://img.shields.io/github/contributors/xinnan-tech/xiaozhi-esp32-server" />
|
||||
<img alt="GitHub Contributors" src="https://img.shields.io/github/contributors/xinnan-tech/xiaozhi-esp32-server?logo=github" />
|
||||
</a>
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/issues">
|
||||
<img alt="Issues" src="https://img.shields.io/github/issues/xinnan-tech/xiaozhi-esp32-server?color=0088ff" />
|
||||
@@ -10,79 +31,114 @@
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/pulls">
|
||||
<img alt="GitHub pull requests" src="https://img.shields.io/github/issues-pr/xinnan-tech/xiaozhi-esp32-server?color=0088ff" />
|
||||
</a>
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server/pulls">
|
||||
<img alt="GitHub pull requests" src="https://img.shields.io/badge/license-MIT-white?labelColor=black" />
|
||||
</a>
|
||||
<a href="https://github.com/xinnan-tech/xiaozhi-esp32-server">
|
||||
<img alt="GitHub pull requests" src="https://img.shields.io/github/stars/xinnan-tech/xiaozhi-esp32-server?color=ffcb47&labelColor=black" />
|
||||
</a>
|
||||
</p>
|
||||
|
||||
# XiaoZhi ESP-32 Backend Service (xiaozhi-esp32-server)
|
||||
|
||||
([中文](README.md) | English)
|
||||
|
||||
This project provides the backend service for the open source smart hardware project [xiaozhi-esp32](https://github.com/78/xiaozhi-esp32). It is implemented in `Python` based on the [XiaoZhi Communication Protocol](https://ccnphfhqs21z.feishu.cn/wiki/M0XiwldO9iJwHikpXD5cEx71nKh).
|
||||
|
||||
---
|
||||
|
||||
## Target Audience 👥
|
||||
## Target Users 👥
|
||||
|
||||
This project is designed to be used in conjunction with ESP32 hardware devices. If you have already purchased an ESP32 device, successfully connected to the backend service deployed by XieGe, and now wish to set up your own `xiaozhi-esp32` backend service, then this project is perfect for you.
|
||||
This project requires ESP32 hardware devices. If you have purchased ESP32-related hardware, successfully connected to Brother Xia's backend service, and wish to set up your own `xiaozhi-esp32` backend service, then this project is perfect for you.
|
||||
|
||||
Want to see it in action? Check out the videos 🎥
|
||||
Want to see it in action? Check out these videos 🎥
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1FMFyejExX" target="_blank">
|
||||
<picture>
|
||||
<img alt="XiaoZhi ESP32 connecting to a custom backend model" src="docs/images/demo1.png" />
|
||||
</picture>
|
||||
</a>
|
||||
<a href="https://www.bilibili.com/video/BV1FMFyejExX" target="_blank">
|
||||
<picture>
|
||||
<img alt="Xiaozhi esp32 connecting to your own backend model" src="docs/images/demo1.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1CDKWemEU6" target="_blank">
|
||||
<picture>
|
||||
<img alt="Custom Voice" src="docs/images/demo2.png" />
|
||||
</picture>
|
||||
</a>
|
||||
<a href="https://www.bilibili.com/video/BV1CDKWemEU6" target="_blank">
|
||||
<picture>
|
||||
<img alt="Custom voice" src="docs/images/demo2.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV12yA2egEaC" target="_blank">
|
||||
<picture>
|
||||
<img alt="Conversing in Cantonese" src="docs/images/demo3.png" />
|
||||
</picture>
|
||||
</a>
|
||||
<a href="https://www.bilibili.com/video/BV12yA2egEaC" target="_blank">
|
||||
<picture>
|
||||
<img alt="Communicating in Cantonese" src="docs/images/demo3.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/av114036381327149" target="_blank">
|
||||
<picture>
|
||||
<img alt="Control Home Appliances" src="docs/images/demo5.png" />
|
||||
</picture>
|
||||
</a>
|
||||
<a href="https://www.bilibili.com/video/BV1pNXWYGEx1" target="_blank">
|
||||
<picture>
|
||||
<img alt="Control home appliances" src="docs/images/demo5.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1kgA2eYEQ9" target="_blank">
|
||||
<picture>
|
||||
<img alt="Lowest Cost Configuration" src="docs/images/demo4.png" />
|
||||
</picture>
|
||||
</a>
|
||||
<a href="https://www.bilibili.com/video/BV1kgA2eYEQ9" target="_blank">
|
||||
<picture>
|
||||
<img alt="Lowest cost configuration" src="docs/images/demo4.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1Vy96YCE3R" target="_blank">
|
||||
<picture>
|
||||
<img alt="Custom voice" src="docs/images/demo6.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1VC96Y5EMH" target="_blank">
|
||||
<picture>
|
||||
<img alt="Play music" src="docs/images/demo7.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV1Z8XuYZEAS" target="_blank">
|
||||
<picture>
|
||||
<img alt="Weather plugin" src="docs/images/demo8.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV178XuYfEpi" target="_blank">
|
||||
<picture>
|
||||
<img alt="IOT command control device" src="docs/images/demo9.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.bilibili.com/video/BV17LXWYvENb" target="_blank">
|
||||
<picture>
|
||||
<img alt="News broadcast" src="docs/images/demo0.png" />
|
||||
</picture>
|
||||
</a>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
---
|
||||
|
||||
## System Requirements and Deployment Prerequisites 🖥️
|
||||
## System Requirements and Prerequisites 🖥️
|
||||
|
||||
- **Hardware**: A set of devices compatible with `xiaozhi-esp32` (for specific models, please refer to [this link](https://rcnv1t9vps13.feishu.cn/wiki/DdgIw4BUgivWDPkhMj1cGIYCnRf)).
|
||||
- **Server**: A computer with at least a 4-core CPU and 8GB of memory.
|
||||
- **Firmware Compilation**: Please update the backend service API endpoint in the `xiaozhi-esp32` project, then recompile the firmware and flash it to your device.
|
||||
- **Computer or Server**: Recommended 4-core CPU, 8GB RAM computer. If using ASR with API, can run on a 2-core CPU, 2GB RAM server.
|
||||
- **Update Client Interface**: Please update the backend service interface address in the client.
|
||||
|
||||
---
|
||||
|
||||
## Warning ⚠️
|
||||
|
||||
This project is relatively new and has not yet undergone network security evaluations. **Do not use it in a production environment.**
|
||||
1. This is open-source software. This software and any third-party API service providers it interfaces with (including but not limited to speech recognition, large language models, speech synthesis, and other platforms) have no commercial partnership. We do not provide any form of guarantee for their service quality or financial security.
|
||||
We recommend users prioritize service providers with relevant business licenses and carefully read their service agreements and privacy policies. This software does not host any account keys, does not participate in fund transfers, and does not bear the risk of recharge fund losses.
|
||||
|
||||
If you deploy this project on a public network for learning purposes, be sure to enable protection in the configuration file `config.yaml`:
|
||||
2. This project is relatively new and has not yet passed network security testing. Please do not use it in production environments. If you deploy this project for learning purposes in a public network environment, please make sure to enable protection in the `config.yaml` configuration file:
|
||||
|
||||
```yaml
|
||||
server:
|
||||
@@ -91,202 +147,150 @@ server:
|
||||
enabled: true
|
||||
```
|
||||
|
||||
Once protection is enabled, you will need to validate the machine's token or MAC address based on your actual situation. Please refer to the configuration documentation for details.
|
||||
After enabling protection, you need to verify the machine's token or MAC address according to actual circumstances. Please refer to the configuration documentation for details.
|
||||
|
||||
---
|
||||
|
||||
## Deployment Methods 🚀
|
||||
|
||||
### I. [Deployment Guide](./docs/Deployment.md)
|
||||
|
||||
This project supports three deployment methods. You can choose based on your actual needs.
|
||||
|
||||
1. [Quick Docker Deployment](./docs/Deployment.md)
|
||||
|
||||
Suitable for regular users who want to quickly experience without much environment configuration. The downside is that pulling the image can be slow. Video tutorial available: [Beautiful expert teaches Docker deployment](https://www.bilibili.com/video/BV1RNQnYDE5t)
|
||||
|
||||
2. [Deploy Using Docker Environment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E5%80%9F%E5%8A%A9docker%E7%8E%AF%E5%A2%83%E8%BF%90%E8%A1%8C%E9%83%A8%E7%BD%B2)
|
||||
|
||||
For software engineers who have Docker installed and want to make custom code modifications.
|
||||
|
||||
3. [Local Source Code Run](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%89%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C)
|
||||
|
||||
Suitable for users familiar with `Conda` environment or who want to build the running environment from scratch.
|
||||
|
||||
For scenarios requiring higher response speed, we recommend using the local source code run method to reduce additional overhead. Video tutorial available: [Handsome expert teaches source code deployment](https://www.bilibili.com/video/BV1GvQWYZEd2)
|
||||
|
||||
### II. [Firmware Compilation](./docs/firmware-build.md)
|
||||
|
||||
Click here to view the detailed process of [firmware compilation](./docs/firmware-build.md).
|
||||
|
||||
After successful flashing and network connection, wake up Xiaozhi using the wake word and pay attention to the console output on the server side.
|
||||
|
||||
---
|
||||
|
||||
## Common Questions ❓
|
||||
|
||||
For issues or product suggestions, please [click here](docs/FAQ.md).
|
||||
|
||||
---
|
||||
|
||||
## Product Ecosystem 👬
|
||||
Xiaozhi is an ecosystem. When using this product, you might want to check out other excellent projects in this ecosystem:
|
||||
|
||||
- [Xiaozhi Android Client](https://github.com/TOM88812/xiaozhi-android-client)
|
||||
A voice dialogue application based on xiaozhi-server for Android and iOS, supporting real-time voice interaction and text dialogue. Now in Flutter version, supporting both iOS and Android.
|
||||
- [Xiaozhi PC Client](https://github.com/Huang-junsen/py-xiaozhi)
|
||||
This project provides a Python-based Xiaobai AI client, allowing you to experience Xiaozhi AI features through code even without physical hardware. Main features include AI voice interaction, visual multimodal recognition, IoT device integration, online music playback, voice wake-up, automatic dialogue mode, graphical interface, command-line mode, cross-platform support, volume control, session management, encrypted audio transmission, automatic verification code processing, etc.
|
||||
- [Xiaozhi Java Server](https://github.com/Huang-junsen/xiaozhi-java)
|
||||
The Xiaozhi open-source backend service Java version is a Java-based open-source project that includes both frontend and backend services, aiming to provide users with a complete backend service solution.
|
||||
---
|
||||
## Feature List ✨
|
||||
|
||||
### Implemented ✅
|
||||
|
||||
- **Communication Protocol**
|
||||
Based on the `xiaozhi-esp32` protocol, data exchange is implemented via WebSocket.
|
||||
Based on `xiaozhi-esp32` protocol, implementing data interaction through WebSocket.
|
||||
- **Dialogue Interaction**
|
||||
Supports wake-up dialogues, manual conversations, and real-time interruptions. Automatically enters sleep mode after long periods of inactivity.
|
||||
- **Multilingual Recognition**
|
||||
Supports Mandarin, Cantonese, English, Japanese, and Korean (default using FunASR).
|
||||
Supports wake-up dialogue, manual dialogue, and real-time interruption. Automatically sleeps after long periods without dialogue
|
||||
- **Intent Recognition**
|
||||
Supports LLM intent recognition and function call, reducing hard-coded intent judgment
|
||||
- **Multi-language Recognition**
|
||||
Supports Mandarin, Cantonese, English, Japanese, Korean (default using FunASR).
|
||||
- **LLM Module**
|
||||
Allows flexible switching of LLM modules. The default is ChatGLMLLM, with options to use AliLLM, DeepSeek, Ollama, and others.
|
||||
Supports flexible switching of LLM modules, default using ChatGLMLLM, can also use Alibaba Bailian, DeepSeek, Ollama, and other interfaces.
|
||||
- **TTS Module**
|
||||
Supports multiple TTS interfaces including EdgeTTS (default) and Volcano Engine Doubao TTS to meet speech synthesis requirements.
|
||||
Supports EdgeTTS (default), Volcano Engine Doubao TTS, and other TTS interfaces to meet speech synthesis needs.
|
||||
- **Memory Function**
|
||||
Supports ultra-long memory, local summary memory, and no memory modes to meet different scenario needs.
|
||||
- **IOT Function**
|
||||
Supports managing registered device IOT functions, supporting intelligent IoT control based on dialogue context.
|
||||
|
||||
### In Development 🚧
|
||||
### Under Development 🚧
|
||||
|
||||
- Conversation Memory Feature
|
||||
- Multiple Mood Modes
|
||||
- Smart Control Panel Web UI
|
||||
- Multiple mood modes
|
||||
- Smart control panel webui
|
||||
|
||||

|
||||
To learn about specific development progress, [click here](https://github.com/users/xinnan-tech/projects/3)
|
||||
|
||||
If you are a software developer, here's an [Open Letter to Developers](docs/contributor_open_letter.md), welcome to join!
|
||||
|
||||
---
|
||||
|
||||
## Supported Platforms/Components 📋
|
||||
## Supported Platforms/Components List 📋
|
||||
|
||||
### LLM
|
||||
### LLM Language Models
|
||||
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Remarks |
|
||||
|:----:|:-----------------------------:|:-----------------------------:|:-----------------:|:-------------------------------------------------------------------------:|
|
||||
| LLM | AliLLM (阿里百炼) | OpenAI API call | Token consumption | [Click to apply for API key](https://bailian.console.aliyun.com/?apiKey=1#/api-key) |
|
||||
| LLM | DeepSeekLLM (深度求索) | OpenAI API call | Token consumption | [Click to apply for API key](https://platform.deepseek.com/) |
|
||||
| LLM | ChatGLMLLM (智谱) | OpenAI API call | Free | Although free, you still need to [click to apply for an API key](https://bigmodel.cn/usercenter/proj-mgmt/apikeys) |
|
||||
| LLM | OllamaLLM | Ollama API call | Free/Custom | Requires pre-downloading the model (`ollama pull`); service URL: `http://localhost:11434` |
|
||||
| LLM | DifyLLM | Dify API call | Token consumption | For local deployment. Note that prompt configuration must be set in the Dify console. |
|
||||
| LLM | GeminiLLM | Gemini API call | Free | [Click to apply for API key](https://aistudio.google.com/apikey) |
|
||||
| LLM | CozeLLM | Coze API call | Token consumption | Requires providing bot_id, user_id, and personal token. |
|
||||
| LLM | Home Assistant | Home Assistant voice assistant API call | Free | Requires providing a Home Assistant token. |
|
||||
| Usage Method | Supported Platforms | Free Platforms |
|
||||
|:---:|:---:|:---:|
|
||||
| openai interface call | Alibaba Bailian, Volcano Engine Doubao, DeepSeek, Zhipu ChatGLM, Gemini | Zhipu ChatGLM, Gemini |
|
||||
| ollama interface call | Ollama | - |
|
||||
| dify interface call | Dify | - |
|
||||
| fastgpt interface call | Fastgpt | - |
|
||||
| coze interface call | Coze | - |
|
||||
|
||||
In fact, any LLM that supports OpenAI API calls can be integrated.
|
||||
In fact, any LLM supporting openai interface calls can be integrated.
|
||||
|
||||
---
|
||||
|
||||
### TTS
|
||||
### TTS Speech Synthesis
|
||||
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Remarks |
|
||||
|:----:|:--------------------------------------:|:------------:|:-----------------:|:--------------------------------------------------------------------------------------:|
|
||||
| TTS | EdgeTTS | API call | Free | Default TTS based on Microsoft's speech synthesis technology. |
|
||||
| TTS | DoubaoTTS (火山引擎豆包 TTS) | API call | Token consumption | [Click to create an API key](https://console.volcengine.com/speech/service/8); it is recommended to use the paid version for higher concurrency. |
|
||||
| TTS | CosyVoiceSiliconflow | API call | Token consumption | Requires application for the Siliconflow API key; output format is WAV. |
|
||||
| TTS | CozeCnTTS | API call | Token consumption | Requires providing a Coze API key; output format is WAV. |
|
||||
| TTS | FishSpeech | API call | Free/Custom | Starts a local TTS service; see the configuration file for startup instructions. |
|
||||
| TTS | GPT_SOVITS_V2 | API call | Free/Custom | Starts a local TTS service, suitable for personalized speech synthesis scenarios. |
|
||||
| Usage Method | Supported Platforms | Free Platforms |
|
||||
|:---:|:---:|:---:|
|
||||
| Interface Call | EdgeTTS, Volcano Engine Doubao TTS, Tencent Cloud, Alibaba Cloud TTS, CosyVoiceSiliconflow, TTS302AI, CozeCnTTS, GizwitsTTS, ACGNTTS, OpenAITTS | EdgeTTS, CosyVoiceSiliconflow(partial) |
|
||||
| Local Service | FishSpeech, GPT_SOVITS_V2, GPT_SOVITS_V3, MinimaxTTS | FishSpeech, GPT_SOVITS_V2, GPT_SOVITS_V3, MinimaxTTS |
|
||||
|
||||
---
|
||||
|
||||
### VAD
|
||||
### VAD Voice Activity Detection
|
||||
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Remarks |
|
||||
|:----:|:-------------------:|:------------:|:-------------:|:-------:|
|
||||
| VAD | SileroVAD | Local | Free | |
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Notes |
|
||||
|:---:|:---------:|:----:|:----:|:--:|
|
||||
| VAD | SileroVAD | Local Use | Free | |
|
||||
|
||||
---
|
||||
|
||||
### ASR
|
||||
### ASR Speech Recognition
|
||||
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Remarks |
|
||||
|:----:|:-------------------:|:------------:|:-------------:|:-------:|
|
||||
| ASR | FunASR | Local | Free | |
|
||||
| ASR | SherpaASR | Local | Free | |
|
||||
| ASR | DoubaoASR | API call | Paid | |
|
||||
| Usage Method | Supported Platforms | Free Platforms |
|
||||
|:---:|:---:|:---:|
|
||||
| Local Use | FunASR, SherpaASR | FunASR, SherpaASR |
|
||||
| Interface Call | DoubaoASR | - |
|
||||
|
||||
---
|
||||
|
||||
## Usage 🚀
|
||||
### Memory Storage
|
||||
|
||||
### 1. [Deployment Documentation](./docs/Deployment.md)
|
||||
|
||||
This project supports three deployment methods. Choose the one that best fits your needs.
|
||||
|
||||
The documentation provided here is a **written tutorial**. If you prefer a **video tutorial**, you can refer to [this expert's hands-on guide](https://www.bilibili.com/video/BV1gePuejEvT).
|
||||
|
||||
Combining both the written and video tutorials can help you get started more quickly.
|
||||
|
||||
1. [Docker Quick Deployment](./docs/Deployment.md)
|
||||
Suitable for general users who want a quick experience without extensive environment configuration. The only downside is that pulling the image can be a bit slow.
|
||||
|
||||
2. [Deployment Using Docker Environment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E5%80%9F%E5%8A%A9docker%E7%8E%AF%E5%A2%83%E8%BF%90%E8%A1%8C%E9%83%A8%E7%BD%B2)
|
||||
Ideal for software engineers who already have Docker installed and wish to customize the code.
|
||||
|
||||
3. [Running from Local Source Code](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%89%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C)
|
||||
Suitable for users familiar with the `Conda` environment or those who wish to build the runtime environment from scratch.
|
||||
|
||||
For scenarios requiring higher response speeds, running from the local source code is recommended to reduce additional overhead.
|
||||
|
||||
### 2. [Firmware Compilation](./docs/firmware-build.md)
|
||||
|
||||
Click [here](./docs/firmware-build.md) for a detailed guide on firmware compilation.
|
||||
|
||||
After successful compilation and network connection, wake up XiaoZhi using the wake-up word and monitor the server console for output.
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Notes |
|
||||
|:------:|:---------------:|:----:|:---------:|:--:|
|
||||
| Memory | mem0ai | Interface Call | 1000 times/month quota | |
|
||||
| Memory | mem_local_short | Local Summary | Free | |
|
||||
|
||||
---
|
||||
|
||||
## Frequently Asked Questions ❓
|
||||
### Intent Recognition
|
||||
|
||||
### 1. TTS often fails and times out ⏰
|
||||
|
||||
**Suggestion:**
|
||||
If `EdgeTTS` frequently fails, please first check whether you are using a proxy (VPN). If so, try disabling the proxy and try again. If you are using Volcano Engine Doubao TTS and it often fails, it is recommended to use the paid version since the trial only supports 2 concurrent requests.
|
||||
|
||||
### 2. I want to control lights, air conditioners, remote power on/off, etc. with XiaoZhi 💡
|
||||
|
||||
**Suggestion:**
|
||||
Set the `LLM` to `HomeAssistant` in the configuration file and use the `HomeAssistant` API to perform the relevant controls.
|
||||
|
||||
### 3. I speak slowly, and XiaoZhi always interrupts during pauses 🗣️
|
||||
|
||||
**Suggestion:**
|
||||
Locate the following section in the configuration file and increase the value of `min_silence_duration_ms` (for example, change it to `1000`):
|
||||
|
||||
```yaml
|
||||
VAD:
|
||||
SileroVAD:
|
||||
threshold: 0.5
|
||||
model_dir: models/snakers4_silero-vad
|
||||
min_silence_duration_ms: 700 # If your pauses are longer, increase this value
|
||||
```
|
||||
|
||||
### 4. Why does XiaoZhi recognize a lot of Korean, Japanese, and English in what I say? 🇰🇷
|
||||
|
||||
**Suggestion:**
|
||||
Check whether the `model.pt` file exists in the `models/SenseVoiceSmall` directory. If it does not, please download it. See [Download ASR Model Files](docs/Deployment.md#模型文件) for details.
|
||||
|
||||
### 5. Why does the error “TTS task error: file does not exist” occur? 📁
|
||||
|
||||
**Suggestion:**
|
||||
Verify that you have correctly installed the `libopus` and `ffmpeg` libraries using `conda`. If not, install them using:
|
||||
|
||||
```
|
||||
conda install conda-forge::libopus
|
||||
conda install conda-forge::ffmpeg
|
||||
```
|
||||
|
||||
### 6. How can I improve XiaoZhi's dialogue response speed? ⚡
|
||||
|
||||
The default configuration of this project is designed to be cost-effective. It is recommended that beginners first use the default free models to ensure that the system runs smoothly, then optimize for faster response times.
|
||||
To improve response speed, you can try replacing individual components. Below are the response time test results for each component (for reference only, not a guarantee):
|
||||
|
||||
**LLM Performance Ranking:**
|
||||
|
||||
| Module Name | Average First Token Time | Average Total Response Time |
|
||||
|--------------|--------------------------|-----------------------------|
|
||||
| AliLLM | 0.547s | 1.485s |
|
||||
| ChatGLMLLM | 0.677s | 3.057s |
|
||||
| OllamaLLM | 0.003s | 0.003s |
|
||||
|
||||
**TTS Performance Ranking:**
|
||||
|
||||
| Module Name | Average Synthesis Time |
|
||||
|----------------------------|------------------------|
|
||||
| EdgeTTS | 1.019s |
|
||||
| DoubaoTTS | 0.503s |
|
||||
| CosyVoiceSiliconflow | 3.732s |
|
||||
|
||||
**Recommended Configuration Combination (Overall Response Speed):**
|
||||
|
||||
| Combination Scheme | Overall Score | LLM First Token | TTS Synthesis |
|
||||
|-----------------------------------|---------------|-----------------|---------------|
|
||||
| AliLLM + DoubaoTTS | 0.539 | 0.547s | 0.503s |
|
||||
| AliLLM + EdgeTTS | 0.642 | 0.547s | 1.019s |
|
||||
| ChatGLMLLM + DoubaoTTS | 0.642 | 0.677s | 0.503s |
|
||||
| ChatGLMLLM + EdgeTTS | 0.745 | 0.677s | 1.019s |
|
||||
| AliLLM + CosyVoiceSiliconflow | 1.184 | 0.547s | 3.732s |
|
||||
|
||||
**Conclusion 🔍**
|
||||
|
||||
_As of February 19, 2025, if my computer were located in Haizhu District, Guangzhou, Guangdong Province, and connected via China Unicom, I would prioritize using:_
|
||||
|
||||
- **LLM:** `AliLLM`
|
||||
- **TTS:** `DoubaoTTS`
|
||||
|
||||
### 7. For more questions, feel free to contact us for feedback 💬
|
||||
|
||||
Our contact information is in [Baidu Netdisk](https://pan.baidu.com/s/1x6USjvP1nTRsZ45XlJu65Q),The extraction code is`223y`。
|
||||
| Type | Platform Name | Usage Method | Pricing Model | Notes |
|
||||
|:------:|:-------------:|:----:|:-------:|:---------------------:|
|
||||
| Intent | intent_llm | Interface Call | Based on LLM pricing | Intent recognition through large models, highly generalizable |
|
||||
| Intent | function_call | Interface Call | Based on LLM pricing | Intent completion through large model function calls, fast and effective |
|
||||
|
||||
---
|
||||
|
||||
## Acknowledgements 🙏
|
||||
## Acknowledgments 🙏
|
||||
|
||||
- This project was inspired by the [Bailing Voice Dialogue Robot](https://github.com/wwbin2017/bailing) and implemented based on it.
|
||||
- Many thanks to [Tenclass](https://www.tenclass.com/) for providing detailed documentation support for the XiaoZhi communication protocol.
|
||||
- This project was inspired by [Bailing Voice Dialogue Robot](https://github.com/wwbin2017/bailing) and implemented based on it.
|
||||
- Thanks to [Tenclass](https://www.tenclass.com/) for providing detailed documentation support for the Xiaozhi communication protocol.
|
||||
|
||||
<a href="https://star-history.com/#xinnan-tech/xiaozhi-esp32-server&Date">
|
||||
<picture>
|
||||
@@ -294,4 +298,4 @@ Our contact information is in [Baidu Netdisk](https://pan.baidu.com/s/1x6USjvP1n
|
||||
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=xinnan-tech/xiaozhi-esp32-server&type=Date" />
|
||||
<img alt="Star History Chart" src="https://api.star-history.com/svg?repos=xinnan-tech/xiaozhi-esp32-server&type=Date" />
|
||||
</picture>
|
||||
</a>
|
||||
</a>
|
||||
Reference in New Issue
Block a user