Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a2023a65c0 | ||
|
|
619370f36c | ||
|
|
ea6c0e691f | ||
|
|
de0d82f1e8 | ||
|
|
c8a26d9c5e | ||
|
|
d89b842586 | ||
|
|
6dd3ceea7e | ||
|
|
5fe68030b0 | ||
|
|
5215105440 | ||
|
|
fb69eb751a | ||
|
|
b903d51592 | ||
|
|
d1f6774e58 | ||
|
|
9e55df4141 | ||
|
|
b169e9413f | ||
|
|
c511766e9d | ||
|
|
e6f9ceb7fa | ||
|
|
eab8b48c9c | ||
|
|
46abb1009a | ||
|
|
2a42d86db0 | ||
|
|
68b0d9654c | ||
|
|
8f96cf13ac | ||
|
|
dbd179e507 | ||
|
|
9541e85600 | ||
|
|
9ecd829bfa | ||
|
|
fd4cff73ae | ||
|
|
1d6c4751de | ||
|
|
e066c1d6a1 | ||
|
|
85f5404b3b | ||
|
|
33c5893d32 | ||
|
|
b597cb0686 | ||
|
|
c5e288ca9c | ||
|
|
38d984f225 | ||
|
|
628c642040 | ||
|
|
d217cc4dd0 | ||
|
|
4f0e54b390 | ||
|
|
d7f89be31c | ||
|
|
53bcee7032 | ||
|
|
96991ae5ef | ||
|
|
068bcde451 | ||
|
|
0dda4f5646 | ||
|
|
6c57ce9dd2 | ||
|
|
6ac67a7e41 | ||
|
|
a23a585ad8 | ||
|
|
eba9562e81 | ||
|
|
ce49b409ac | ||
|
|
7222f68d4d | ||
|
|
f3f0d62f12 | ||
|
|
6e7c86e159 | ||
|
|
f5565f6700 | ||
|
|
e8d0bb0c54 | ||
|
|
33d70ccc96 | ||
|
|
53313a26af | ||
|
|
7ba180752a | ||
|
|
19736e66ad | ||
|
|
a00f8e4b76 | ||
|
|
7d9895cf5b | ||
|
|
d5f804bbb3 | ||
|
|
33a385cfa8 | ||
|
|
109d924591 | ||
|
|
a825eb3d4c | ||
|
|
3fb40677a4 | ||
|
|
833971cd28 | ||
|
|
d14b14bce9 | ||
|
|
53e26821ad | ||
|
|
1a7c06eb81 | ||
|
|
42a5b4892d | ||
|
|
9f4508b0c7 | ||
|
|
b2123ff01a | ||
|
|
43ead841a4 | ||
|
|
33b4794e83 | ||
|
|
fb91e6b1dd | ||
|
|
4a4dbf123e | ||
|
|
0decedd6a1 | ||
|
|
d2e3a63418 | ||
|
|
5eaaf9f01d | ||
|
|
d6697948c2 | ||
|
|
80d10cf4b5 | ||
|
|
ba1fb16f2d | ||
|
|
5f574667d3 | ||
|
|
c470bb1db1 | ||
|
|
26ed8ca33f | ||
|
|
3c46e16494 |
@@ -183,8 +183,8 @@ Spearheaded by Professor Siyuan Liu's Team (South China University of Technology
|
||||
#### 🚀 部署方式选择
|
||||
| 部署方式 | 特点 | 适用场景 | 部署文档 | 配置要求 | 视频教程 |
|
||||
|---------|------|---------|---------|---------|---------|
|
||||
| **最简化安装** | 智能对话、IOT、MCP、视觉感知 | 低配置环境,数据存储在配置文件,无需数据库 | [①Docker版](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②源码部署](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 如果使用`FunASR`要2核4G,如果全API,要2核2G | - |
|
||||
| **全模块安装** | 智能对话、IOT、MCP接入点、声纹识别、视觉感知、OTA、智控台 | 完整功能体验,数据存储在数据库 |[①Docker版](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②源码部署](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③源码部署自动更新教程](./docs/dev-ops-integration.md) | 如果使用`FunASR`要4核8G,如果全API,要2核4G| [本地源码启动视频教程](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
| **最简化安装** | 智能对话、单智能体管理 | 低配置环境,数据存储在配置文件,无需数据库 | [①Docker版](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②源码部署](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 如果使用`FunASR`要2核4G,如果全API,要2核2G | - |
|
||||
| **全模块安装** | 智能对话、多用户管理、多智能体管理、智控台界面操作 | 完整功能体验,数据存储在数据库 |[①Docker版](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②源码部署](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③源码部署自动更新教程](./docs/dev-ops-integration.md) | 如果使用`FunASR`要4核8G,如果全API,要2核4G| [本地源码启动视频教程](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
|
||||
常见问题及相关教程,可参考[这个链接](./docs/FAQ.md)
|
||||
|
||||
@@ -211,10 +211,10 @@ Websocket接口地址: wss://2662r3426b.vicp.fun/xiaozhi/v1/
|
||||
|
||||
| 模块名称 | 入门全免费设置 | 流式配置 |
|
||||
|:---:|:---:|:---:|
|
||||
| ASR(语音识别) | FunASR(本地) | 👍FunASR(本地GPU模式) |
|
||||
| LLM(大模型) | ChatGLMLLM(智谱glm-4-flash) | 👍AliLLM(qwen3-235b-a22b-instruct-2507) 或 👍DoubaoLLM(doubao-1-5-pro-32k-250115) |
|
||||
| VLLM(视觉大模型) | ChatGLMVLLM(智谱glm-4v-flash) | 👍QwenVLVLLM(千问qwen2.5-vl-3b-instructh) |
|
||||
| TTS(语音合成) | ✅LinkeraiTTS(灵犀流式) | 👍HuoshanDoubleStreamTTS(火山双流式语音合成) 或 👍AliyunStreamTTS(阿里云流式语音合成) |
|
||||
| ASR(语音识别) | FunASR(本地) | 👍XunfeiStreamASR(讯飞流式) |
|
||||
| LLM(大模型) | glm-4-flash(智谱) | 👍qwen-flash(阿里百炼) |
|
||||
| VLLM(视觉大模型) | glm-4v-flash(智谱) | 👍qwen2.5-vl-3b-instructh(阿里百炼) |
|
||||
| TTS(语音合成) | ✅LinkeraiTTS(灵犀流式) | 👍HuoshanDoubleStreamTTS(火山流式) |
|
||||
| Intent(意图识别) | function_call(函数调用) | function_call(函数调用) |
|
||||
| Memory(记忆功能) | mem_local_short(本地短期记忆) | mem_local_short(本地短期记忆) |
|
||||
|
||||
@@ -260,7 +260,7 @@ Websocket接口地址: wss://2662r3426b.vicp.fun/xiaozhi/v1/
|
||||
---
|
||||
|
||||
## 产品生态 👬
|
||||
小智是一个生态,当你使用这个产品时,也可以看看其他在这个生态圈的[优秀项目](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#%E7%9B%B8%E5%85%B3%E5%BC%80%E6%BA%90%E9%A1%B9%E7%9B%AE)
|
||||
小智是一个生态,当你使用这个产品时,也可以看看其他在这个生态圈的[优秀项目](https://github.com/78/xiaozhi-esp32/blob/main/README_zh.md#%E7%9B%B8%E5%85%B3%E5%BC%80%E6%BA%90%E9%A1%B9%E7%9B%AE)
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -181,8 +181,8 @@ Dieses Projekt bietet zwei Bereitstellungsmethoden. Bitte wählen Sie basierend
|
||||
#### 🚀 Auswahl der Bereitstellungsmethode
|
||||
| Bereitstellungsmethode | Funktionen | Anwendungsszenarien | Deployment-Dokumente | Konfigurationsanforderungen | Video-Tutorials |
|
||||
|---------|------|---------|---------|---------|---------|
|
||||
| **Vereinfachte Installation** | Intelligenter Dialog, IOT, MCP, visuelle Wahrnehmung | Umgebungen mit geringer Konfiguration, Daten in Konfigurationsdateien gespeichert, keine Datenbank erforderlich | [①Docker-Version](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Quellcode-Deployment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 Kerne 4GB bei Verwendung von `FunASR`, 2 Kerne 2GB bei allen APIs | - |
|
||||
| **Vollständige Modulinstallation** | Intelligenter Dialog, IOT, MCP-Endpunkte, Stimmabdruckerkennung, visuelle Wahrnehmung, OTA, intelligente Steuerkonsole | Vollständige Funktionserfahrung, Daten in Datenbank gespeichert |[①Docker-Version](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Quellcode-Deployment](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Quellcode-Deployment Auto-Update-Tutorial](./docs/dev-ops-integration.md) | 4 Kerne 8GB bei Verwendung von `FunASR`, 2 Kerne 4GB bei allen APIs| [Video-Tutorial für lokalen Quellcode-Start](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
| **Vereinfachte Installation** | Intelligenter Dialog, Einzel-Agenten-Verwaltung | Umgebungen mit geringer Konfiguration, Daten in Konfigurationsdateien gespeichert, keine Datenbank erforderlich | [①Docker-Version](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Quellcode-Deployment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 Kerne 4GB bei Verwendung von `FunASR`, 2 Kerne 2GB bei allen APIs | - |
|
||||
| **Vollständige Modulinstallation** | Intelligenter Dialog, Mehrbenutzerverwaltung, Mehr-Agenten-Verwaltung, Intelligente Steuerkonsole-Bedienung | Vollständige Funktionserfahrung, Daten in Datenbank gespeichert |[①Docker-Version](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Quellcode-Deployment](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Quellcode-Deployment Auto-Update-Tutorial](./docs/dev-ops-integration.md) | 4 Kerne 8GB bei Verwendung von `FunASR`, 2 Kerne 4GB bei allen APIs| [Video-Tutorial für lokalen Quellcode-Start](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
|
||||
Häufige Fragen und entsprechende Tutorials finden Sie unter [diesem Link](./docs/FAQ.md)
|
||||
|
||||
@@ -209,10 +209,10 @@ Websocket-Schnittstellenadresse: wss://2662r3426b.vicp.fun/xiaozhi/v1/
|
||||
|
||||
| Modulname | Einstiegslevel Kostenlose Einstellungen | Streaming-Konfiguration |
|
||||
|:---:|:---:|:---:|
|
||||
| ASR (Spracherkennung) | FunASR (Lokal) | 👍FunASR (Lokaler GPU-Modus) |
|
||||
| LLM (Großes Modell) | ChatGLMLLM (Zhipu glm-4-flash) | 👍AliLLM (qwen3-235b-a22b-instruct-2507) oder 👍DoubaoLLM (doubao-1-5-pro-32k-250115) |
|
||||
| VLLM (Vision Large Model) | ChatGLMVLLM (Zhipu glm-4v-flash) | 👍QwenVLVLLM (Qwen qwen2.5-vl-3b-instructh) |
|
||||
| TTS (Sprachsynthese) | ✅LinkeraiTTS (Lingxi-Streaming) | 👍HuoshanDoubleStreamTTS (Volcano Dual-Stream-Sprachsynthese) oder 👍AliyunStreamTTS (Alibaba Cloud Streaming-Sprachsynthese) |
|
||||
| ASR (Spracherkennung) | FunASR (Lokal) | 👍XunfeiStreamASR (Xunfei-Streaming) |
|
||||
| LLM (Großes Modell) | glm-4-flash (Zhipu) | 👍qwen-flash (Alibaba Bailian) |
|
||||
| VLLM (Vision Large Model) | glm-4v-flash (Zhipu) | 👍qwen2.5-vl-3b-instructh (Alibaba Bailian) |
|
||||
| TTS (Sprachsynthese) | ✅LinkeraiTTS (Lingxi-Streaming) | 👍HuoshanDoubleStreamTTS (Volcano-Streaming) |
|
||||
| Intent (Absichtserkennung) | function_call (Funktionsaufruf) | function_call (Funktionsaufruf) |
|
||||
| Memory (Gedächtnisfunktion) | mem_local_short (Lokales Kurzzeitgedächtnis) | mem_local_short (Lokales Kurzzeitgedächtnis) |
|
||||
|
||||
@@ -258,7 +258,7 @@ Wenn Sie ein Softwareentwickler sind, finden Sie hier einen [Offenen Brief an En
|
||||
---
|
||||
|
||||
## Produktökosystem 👬
|
||||
Xiaozhi ist ein Ökosystem. Wenn Sie dieses Produkt verwenden, können Sie sich auch andere [hervorragende Projekte](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#%E7%9B%B8%E5%85%B3%E5%BC%80%E6%BA%90%E9%A1%B9%E7%9B%AE) in diesem Ökosystem ansehen
|
||||
Xiaozhi ist ein Ökosystem. Wenn Sie dieses Produkt verwenden, können Sie sich auch andere [hervorragende Projekte](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#related-open-source-projects) in diesem Ökosystem ansehen
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -181,8 +181,8 @@ This project provides two deployment methods. Please choose based on your specif
|
||||
#### 🚀 Deployment Method Selection
|
||||
| Deployment Method | Features | Applicable Scenarios | Deployment Docs | Configuration Requirements | Video Tutorials |
|
||||
|---------|------|---------|---------|---------|---------|
|
||||
| **Simplified Installation** | Intelligent dialogue, IOT, MCP, visual perception | Low-configuration environments, data stored in config files, no database required | [①Docker Version](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Source Code Deployment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 cores 4GB if using `FunASR`, 2 cores 2GB if all APIs | - |
|
||||
| **Full Module Installation** | Intelligent dialogue, IOT, MCP endpoints, voiceprint recognition, visual perception, OTA, intelligent control console | Complete functionality experience, data stored in database |[①Docker Version](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Source Code Deployment](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Source Code Deployment Auto-Update Tutorial](./docs/dev-ops-integration.md) | 4 cores 8GB if using `FunASR`, 2 cores 4GB if all APIs| [Local Source Code Startup Video Tutorial](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
| **Simplified Installation** | Intelligent dialogue, single agent management | Low-configuration environments, data stored in config files, no database required | [①Docker Version](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Source Code Deployment](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 cores 4GB if using `FunASR`, 2 cores 2GB if all APIs | - |
|
||||
| **Full Module Installation** | Intelligent dialogue, multi-user management, multi-agent management, intelligent console interface operation | Complete functionality experience, data stored in database |[①Docker Version](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Source Code Deployment](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Source Code Deployment Auto-Update Tutorial](./docs/dev-ops-integration.md) | 4 cores 8GB if using `FunASR`, 2 cores 4GB if all APIs| [Local Source Code Startup Video Tutorial](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
|
||||
|
||||
> 💡 Note: Below is a test platform deployed with the latest code. You can burn and test if needed. Concurrent users: 6, data will be cleared daily.
|
||||
@@ -208,10 +208,10 @@ Websocket Interface Address: wss://2662r3426b.vicp.fun/xiaozhi/v1/
|
||||
|
||||
| Module Name | Entry Level Free Settings | Streaming Configuration |
|
||||
|:---:|:---:|:---:|
|
||||
| ASR(Speech Recognition) | FunASR(Local) | 👍FunASRServer or 👍DoubaoStreamASR |
|
||||
| LLM(Large Model) | ChatGLMLLM(Zhipu glm-4-flash) | 👍DoubaoLLM(Volcano doubao-1-5-pro-32k-250115) |
|
||||
| VLLM(Vision Large Model) | ChatGLMVLLM(Zhipu glm-4v-flash) | 👍QwenVLVLLM(Qwen qwen2.5-vl-3b-instructh) |
|
||||
| TTS(Speech Synthesis) | ✅LinkeraiTTS(Lingxi streaming) | 👍HuoshanDoubleStreamTTS(Volcano dual-stream speech synthesis) |
|
||||
| ASR(Speech Recognition) | FunASR(Local) | 👍XunfeiStreamASR(Xunfei Streaming) |
|
||||
| LLM(Large Model) | glm-4-flash(Zhipu) | 👍qwen-flash(Alibaba Bailian) |
|
||||
| VLLM(Vision Large Model) | glm-4v-flash(Zhipu) | 👍qwen2.5-vl-3b-instructh(Alibaba Bailian) |
|
||||
| TTS(Speech Synthesis) | ✅LinkeraiTTS(Lingxi streaming) | 👍HuoshanDoubleStreamTTS(Volcano Streaming) |
|
||||
| Intent(Intent Recognition) | function_call(Function calling) | function_call(Function calling) |
|
||||
| Memory(Memory function) | mem_local_short(Local short-term memory) | mem_local_short(Local short-term memory) |
|
||||
|
||||
@@ -256,7 +256,7 @@ If you are a software developer, here is an [Open Letter to Developers](docs/con
|
||||
---
|
||||
|
||||
## Product Ecosystem 👬
|
||||
Xiaozhi is an ecosystem. When using this product, you can also check out other [excellent projects](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#%E7%9B%B8%E5%85%B3%E5%BC%80%E6%BA%90%E9%A1%B9%E7%9B%AE) in this ecosystem
|
||||
Xiaozhi is an ecosystem. When using this product, you can also check out other [excellent projects](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#related-open-source-projects) in this ecosystem
|
||||
|
||||
| Project Name | Project Address | Project Description |
|
||||
|:---------------------|:--------|:--------|
|
||||
|
||||
@@ -182,8 +182,8 @@ Dự án này cung cấp hai phương pháp triển khai, vui lòng chọn theo
|
||||
#### 🚀 Lựa chọn phương pháp triển khai
|
||||
| Phương pháp triển khai | Đặc điểm | Tình huống áp dụng | Tài liệu triển khai | Yêu cầu cấu hình | Video hướng dẫn |
|
||||
|---------|------|---------|---------|---------|---------|
|
||||
| **Cài đặt tối giản** | Đối thoại thông minh, IOT, MCP, cảm nhận thị giác | Môi trường cấu hình thấp, dữ liệu lưu trong tệp cấu hình, không cần cơ sở dữ liệu | [①Phiên bản Docker](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Triển khai mã nguồn](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 nhân 4GB nếu dùng `FunASR`, 2 nhân 2GB nếu toàn API | - |
|
||||
| **Cài đặt toàn bộ module** | Đối thoại thông minh, IOT, điểm truy cập MCP, nhận dạng giọng nói, cảm nhận thị giác, OTA, bảng điều khiển thông minh | Trải nghiệm đầy đủ tính năng, dữ liệu lưu trong cơ sở dữ liệu |[①Phiên bản Docker](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Triển khai mã nguồn](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Hướng dẫn tự động cập nhật triển khai mã nguồn](./docs/dev-ops-integration.md) | 4 nhân 8GB nếu dùng `FunASR`, 2 nhân 4GB nếu toàn API| [Video hướng dẫn khởi động mã nguồn cục bộ](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
| **Cài đặt tối giản** | Đối thoại thông minh, quản lý đơn tác nhân | Môi trường cấu hình thấp, dữ liệu lưu trong tệp cấu hình, không cần cơ sở dữ liệu | [①Phiên bản Docker](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E5%8F%AA%E8%BF%90%E8%A1%8Cserver) / [②Triển khai mã nguồn](./docs/Deployment.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E5%8F%AA%E8%BF%90%E8%A1%8Cserver)| 2 nhân 4GB nếu dùng `FunASR`, 2 nhân 2GB nếu toàn API | - |
|
||||
| **Cài đặt toàn bộ module** | Đối thoại thông minh, quản lý đa người dùng, quản lý đa tác nhân, bảng điều khiển thông minh | Trải nghiệm đầy đủ tính năng, dữ liệu lưu trong cơ sở dữ liệu |[①Phiên bản Docker](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%B8%80docker%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [②Triển khai mã nguồn](./docs/Deployment_all.md#%E6%96%B9%E5%BC%8F%E4%BA%8C%E6%9C%AC%E5%9C%B0%E6%BA%90%E7%A0%81%E8%BF%90%E8%A1%8C%E5%85%A8%E6%A8%A1%E5%9D%97) / [③Hướng dẫn tự động cập nhật triển khai mã nguồn](./docs/dev-ops-integration.md) | 4 nhân 8GB nếu dùng `FunASR`, 2 nhân 4GB nếu toàn API| [Video hướng dẫn khởi động mã nguồn cục bộ](https://www.bilibili.com/video/BV1wBJhz4Ewe) |
|
||||
|
||||
Câu hỏi thường gặp và hướng dẫn liên quan, vui lòng tham khảo [liên kết này](./docs/FAQ.md)
|
||||
|
||||
@@ -210,10 +210,10 @@ Công cụ kiểm tra dịch vụ: https://2662r3426b.vicp.fun/test/
|
||||
|
||||
| Tên module | Cài đặt miễn phí cho người mới | Cấu hình streaming |
|
||||
|:---:|:---:|:---:|
|
||||
| ASR(Nhận dạng giọng nói) | FunASR(Local) | 👍FunASR(Chế độ GPU cục bộ) |
|
||||
| LLM(Mô hình lớn) | ChatGLMLLM(Zhipu glm-4-flash) | 👍AliLLM(qwen3-235b-a22b-instruct-2507) hoặc 👍DoubaoLLM(doubao-1-5-pro-32k-250115) |
|
||||
| VLLM(Mô hình lớn thị giác) | ChatGLMVLLM(Zhipu glm-4v-flash) | 👍QwenVLVLLM(Qwen qwen2.5-vl-3b-instructh) |
|
||||
| TTS(Tổng hợp giọng nói) | ✅LinkeraiTTS(Lingxi streaming) | 👍HuoshanDoubleStreamTTS(Tổng hợp giọng nói streaming kép Volcano) hoặc 👍AliyunStreamTTS(Tổng hợp giọng nói streaming Alibaba Cloud) |
|
||||
| ASR(Nhận dạng giọng nói) | FunASR(Local) | 👍XunfeiStreamASR(Xunfei Streaming) |
|
||||
| LLM(Mô hình lớn) | glm-4-flash(Zhipu) | 👍qwen-flash(Alibaba Bailian) |
|
||||
| VLLM(Mô hình lớn thị giác) | glm-4v-flash(Zhipu) | 👍qwen2.5-vl-3b-instructh(Alibaba Bailian) |
|
||||
| TTS(Tổng hợp giọng nói) | ✅LinkeraiTTS(Lingxi streaming) | 👍HuoshanDoubleStreamTTS(Volcano Streaming) |
|
||||
| Intent(Nhận dạng ý định) | function_call(Gọi hàm) | function_call(Gọi hàm) |
|
||||
| Memory(Chức năng bộ nhớ) | mem_local_short(Bộ nhớ ngắn hạn cục bộ) | mem_local_short(Bộ nhớ ngắn hạn cục bộ) |
|
||||
|
||||
@@ -259,7 +259,7 @@ Nếu bạn là một nhà phát triển phần mềm, đây có một [Lá thư
|
||||
---
|
||||
|
||||
## Hệ sinh thái sản phẩm 👬
|
||||
Xiaozhi là một hệ sinh thái, khi bạn sử dụng sản phẩm này, bạn cũng có thể xem các [dự án xuất sắc](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#%E7%9B%B8%E5%85%B3%E5%BC%80%E6%BA%90%E9%A1%B9%E7%9B%AE) khác trong hệ sinh thái này
|
||||
Xiaozhi là một hệ sinh thái, khi bạn sử dụng sản phẩm này, bạn cũng có thể xem các [dự án xuất sắc](https://github.com/78/xiaozhi-esp32?tab=readme-ov-file#related-open-source-projects) khác trong hệ sinh thái này
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -38,10 +38,10 @@ conda install conda-forge::ffmpeg
|
||||
|
||||
| 模块名称 | 入门全免费设置 | 流式配置 |
|
||||
|:---:|:---:|:---:|
|
||||
| ASR(语音识别) | FunASR(本地) | 👍FunASR(本地GPU模式) |
|
||||
| LLM(大模型) | ChatGLMLLM(智谱glm-4-flash) | 👍AliLLM(qwen3-235b-a22b-instruct-2507) 或 👍DoubaoLLM(doubao-1-5-pro-32k-250115) |
|
||||
| VLLM(视觉大模型) | ChatGLMVLLM(智谱glm-4v-flash) | 👍QwenVLVLLM(千问qwen2.5-vl-3b-instructh) |
|
||||
| TTS(语音合成) | ✅LinkeraiTTS(灵犀流式) | 👍HuoshanDoubleStreamTTS(火山双流式语音合成) 或 👍AliyunStreamTTS(阿里云流式语音合成) |
|
||||
| ASR(语音识别) | FunASR(本地) | 👍XunfeiStreamASR(讯飞流式) |
|
||||
| LLM(大模型) | glm-4-flash(智谱) | 👍qwen-flash(阿里百炼) |
|
||||
| VLLM(视觉大模型) | glm-4v-flash(智谱) | 👍qwen2.5-vl-3b-instructh(阿里百炼) |
|
||||
| TTS(语音合成) | ✅LinkeraiTTS(灵犀流式) | 👍HuoshanDoubleStreamTTS(火山流式) |
|
||||
| Intent(意图识别) | function_call(函数调用) | function_call(函数调用) |
|
||||
| Memory(记忆功能) | mem_local_short(本地短期记忆) | mem_local_short(本地短期记忆) |
|
||||
|
||||
@@ -69,6 +69,7 @@ VAD:
|
||||
### 9、编译固件相关教程
|
||||
1、[如何自己编译小智固件](./firmware-build.md)<br/>
|
||||
2、[如何基于虾哥编译好的固件修改OTA地址](./firmware-setting.md)<br/>
|
||||
3、[单模块部署如何配置固件OTA自动升级](./ota-upgrade-guide.md)<br/>
|
||||
|
||||
### 10、拓展相关教程
|
||||
1、[如何开启手机号码注册智控台](./ali-sms-integration.md)<br/>
|
||||
|
||||
@@ -23,6 +23,8 @@
|
||||
|
||||
### 2.将音色资源ID分配给系统账号
|
||||
|
||||
使用超级管理员账号登录智控台,点击顶部`参数字典`,在下拉菜单中,点击`系统功能配置`页面。在页面上勾选`音色克隆`,点击保存配置。即可在顶部菜单看到`音色克隆`按钮。
|
||||
|
||||
使用超级管理员账号登录智控台,点击顶部【音色克隆】、【音色资源】。
|
||||
|
||||
点击新增按钮,在【平台名称】选择“火山双流式语音合成”;
|
||||
|
||||
@@ -76,6 +76,7 @@ MQTT_PORT=1883 # MQTT服务器端口
|
||||
UDP_PORT=8884 # UDP服务器端口
|
||||
API_PORT=8007 # 管理API端口
|
||||
MQTT_SIGNATURE_KEY=test # MQTT签名密钥
|
||||
SERVER_SECRET=Te1st12134 # 服务器密钥,请保持和智控台(server.secret)一致或者和xiaozhi-server里(server.auth_key)保持一致
|
||||
```
|
||||
请注意`PUBLIC_IP`配置,确保其与实际公网IP一致,如果有域名就填域名。
|
||||
|
||||
@@ -85,6 +86,13 @@ MQTT_SIGNATURE_KEY=test # MQTT签名密钥
|
||||
- 注意不要用简单的密码,比如`123456`、`test`等。
|
||||
- 注意不要用简单的密码,比如`123456`、`test`等。
|
||||
|
||||
`SERVER_SECRET` 是用生成websocket连接的认证信息。
|
||||
|
||||
1、如果你是全模块部署,且你的智控台的参数管理里`server.auth.enabled`设置成了`true`,那么,`SERVER_SECRET`需要和智控台(`server.secret`)保持一致。
|
||||
|
||||
2、如果你是单模块部署,且你在配置文件里把`server.auth.enabled`设置成了`true`,那么,`SERVER_SECRET`需要和配置文件里(`server.auth_key`)保持一致。
|
||||
|
||||
|
||||
6. 启动MQTT网关
|
||||
```
|
||||
# 启动服务
|
||||
|
||||
@@ -0,0 +1,142 @@
|
||||
# 单模块部署固件OTA自动升级配置指南
|
||||
|
||||
本教程将指导你如何在**单模块部署**场景下配置固件OTA自动升级功能,实现设备固件的自动更新。
|
||||
|
||||
如果你已经使用**全模块部署**,请忽略本教程。
|
||||
|
||||
## 功能介绍
|
||||
|
||||
在单模块部署中,xiaozhi-server内置了OTA固件管理功能,可以自动检测设备版本并下发升级固件。系统会根据设备型号和当前版本,自动匹配并推送最新的固件版本。
|
||||
|
||||
## 前提条件
|
||||
|
||||
- 你已经成功进行**单模块部署**并运行xiaozhi-server
|
||||
- 设备能够正常连接到服务器
|
||||
|
||||
## 第一步 准备固件文件
|
||||
|
||||
### 1. 创建固件存放目录
|
||||
|
||||
固件文件需要放在`data/bin/`目录下。如果该目录不存在,请手动创建:
|
||||
|
||||
```bash
|
||||
mkdir -p data/bin
|
||||
```
|
||||
|
||||
### 2. 固件文件命名规则
|
||||
|
||||
固件文件必须遵循以下命名格式:
|
||||
|
||||
```
|
||||
{设备型号}_{版本号}.bin
|
||||
```
|
||||
|
||||
**命名规则说明:**
|
||||
- `设备型号`:设备的型号名称,例如 `lichuang-dev`、`bread-compact-wifi` 等
|
||||
- `版本号`:固件版本号,必须以数字开头,支持数字、字母、点号、下划线和短横线,例如 `1.6.6`、`2.0.0` 等
|
||||
- 文件扩展名必须是 `.bin`
|
||||
|
||||
**命名示例:**
|
||||
```
|
||||
bread-compact-wifi_1.6.6.bin
|
||||
lichuang-dev_2.0.0.bin
|
||||
```
|
||||
|
||||
### 3. 放置固件文件
|
||||
|
||||
将准备好的固件文件(.bin文件)复制到`data/bin/`目录下:
|
||||
|
||||
重要的事情说三遍:升级的bin文件是`xiaozhi.bin`,不是全量固件文件`merged-binary.bin`!
|
||||
|
||||
重要的事情说三遍:升级的bin文件是`xiaozhi.bin`,不是全量固件文件`merged-binary.bin`!
|
||||
|
||||
重要的事情说三遍:升级的bin文件是`xiaozhi.bin`,不是全量固件文件`merged-binary.bin`!
|
||||
|
||||
```bash
|
||||
cp xiaozhi.bin data/bin/设备型号_版本号.bin
|
||||
```
|
||||
|
||||
例如:
|
||||
```bash
|
||||
cp xiaozhi.bin data/bin/bread-compact-wifi_1.6.6.bin
|
||||
```
|
||||
|
||||
## 第二步 配置公网访问地址(仅公网部署需要)
|
||||
|
||||
**注意:此步骤仅适用于单模块公网部署的场景。**
|
||||
|
||||
如果你的xiaozhi-server是公网部署(使用公网IP或域名),**必须**配置`server.vision_explain`参数,因为OTA固件下载地址会使用该配置的域名和端口。
|
||||
|
||||
如果你是局域网部署,可以跳过此步骤。
|
||||
|
||||
### 为什么要配置这个参数?
|
||||
|
||||
在单模块部署中,系统生成固件下载地址时,会使用`vision_explain`配置的域名和端口作为基础地址。如果不配置或配置错误,设备将无法访问固件下载地址。
|
||||
|
||||
### 配置方法
|
||||
|
||||
打开`data/.config.yaml`文件,找到`server`配置段,设置`vision_explain`参数:
|
||||
|
||||
```yaml
|
||||
server:
|
||||
vision_explain: http://你的域名或IP:端口号/mcp/vision/explain
|
||||
```
|
||||
|
||||
**配置示例:**
|
||||
|
||||
局域网部署(默认):
|
||||
```yaml
|
||||
server:
|
||||
vision_explain: http://192.168.1.100:8003/mcp/vision/explain
|
||||
```
|
||||
|
||||
公网域名部署:
|
||||
```yaml
|
||||
server:
|
||||
vision_explain: http://yourdomain.com:8003/mcp/vision/explain
|
||||
```
|
||||
|
||||
### 注意事项
|
||||
|
||||
- 域名或IP必须是设备能够访问的地址
|
||||
- 如果使用Docker部署,不能使用Docker内部地址(如127.0.0.1或localhost)
|
||||
- 如果你使用了nginx反向代理,请填写对外的地址和端口号,不是本项目运行的端口号
|
||||
|
||||
|
||||
## 常见问题
|
||||
|
||||
### 1. 设备收不到固件更新
|
||||
|
||||
**可能原因和解决方法:**
|
||||
|
||||
- 检查固件文件命名是否符合规则:`{型号}_{版本号}.bin`
|
||||
- 检查固件文件是否正确放置在`data/bin/`目录
|
||||
- 检查设备型号是否与固件文件名中的型号匹配
|
||||
- 检查固件版本号是否高于设备当前版本
|
||||
- 查看服务器日志,确认OTA请求是否正常处理
|
||||
|
||||
### 2. 设备报告下载地址无法访问
|
||||
|
||||
**可能原因和解决方法:**
|
||||
|
||||
- 检查`server.vision_explain`配置的域名或IP是否正确
|
||||
- 确认端口号配置正确(默认8003)
|
||||
- 如果是公网部署,确保设备能够访问该公网地址
|
||||
- 如果是Docker部署,确保不是使用了内部地址(127.0.0.1)
|
||||
- 检查防火墙是否开放了对应端口
|
||||
- 如果你使用了nginx反向代理,请填写对外的地址和端口号,不是本项目运行的端口号
|
||||
|
||||
### 3. 如何确认设备当前版本
|
||||
|
||||
查看OTA请求日志,日志中会显示设备上报的版本号:
|
||||
|
||||
```
|
||||
[ota_handler] - 设备 AA:BB:CC:DD:EE:FF 固件已是最新: 1.6.6
|
||||
```
|
||||
|
||||
### 4. 固件文件放置后没有生效
|
||||
|
||||
系统有30秒的缓存时间(默认),可以:
|
||||
- 等待30秒后再让设备发起OTA请求
|
||||
- 重启xiaozhi-server服务
|
||||
- 调整`firmware_cache_ttl`配置为更短的时间
|
||||
@@ -304,7 +304,7 @@ public interface Constant {
|
||||
/**
|
||||
* 版本号
|
||||
*/
|
||||
public static final String VERSION = "0.8.10";
|
||||
public static final String VERSION = "0.8.11";
|
||||
|
||||
/**
|
||||
* 无效固件URL
|
||||
|
||||
@@ -159,4 +159,11 @@ public class RedisKeys {
|
||||
public static String getKnowledgeBaseCacheKey(String datasetId) {
|
||||
return "knowledge:base:" + datasetId;
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取临时注册设备标记key
|
||||
*/
|
||||
public static String getTmpRegisterMacKey(String deviceId) {
|
||||
return "tmp_register_mac:" + deviceId;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -44,6 +44,7 @@ import xiaozhi.modules.agent.entity.AgentEntity;
|
||||
import xiaozhi.modules.agent.entity.AgentTemplateEntity;
|
||||
import xiaozhi.modules.agent.service.AgentChatAudioService;
|
||||
import xiaozhi.modules.agent.service.AgentChatHistoryService;
|
||||
import xiaozhi.modules.agent.service.AgentChatSummaryService;
|
||||
import xiaozhi.modules.agent.service.AgentContextProviderService;
|
||||
import xiaozhi.modules.agent.service.AgentPluginMappingService;
|
||||
import xiaozhi.modules.agent.service.AgentService;
|
||||
@@ -66,6 +67,7 @@ public class AgentController {
|
||||
private final AgentChatAudioService agentChatAudioService;
|
||||
private final AgentPluginMappingService agentPluginMappingService;
|
||||
private final AgentContextProviderService agentContextProviderService;
|
||||
private final AgentChatSummaryService agentChatSummaryService;
|
||||
private final RedisUtils redisUtils;
|
||||
|
||||
@GetMapping("/list")
|
||||
@@ -119,6 +121,27 @@ public class AgentController {
|
||||
return new Result<>();
|
||||
}
|
||||
|
||||
@PostMapping("/chat-summary/{sessionId}/save")
|
||||
@Operation(summary = "根据会话ID生成聊天记录总结并保存(异步执行)")
|
||||
public Result<Void> generateAndSaveChatSummary(@PathVariable String sessionId) {
|
||||
try {
|
||||
// 异步执行总结生成任务,立即返回成功响应
|
||||
new Thread(() -> {
|
||||
try {
|
||||
agentChatSummaryService.generateAndSaveChatSummary(sessionId);
|
||||
System.out.println("异步执行会话 " + sessionId + " 的聊天记录总结完成");
|
||||
} catch (Exception e) {
|
||||
System.err.println("异步执行会话 " + sessionId + " 的聊天记录总结失败: " + e.getMessage());
|
||||
}
|
||||
}).start();
|
||||
|
||||
// 立即返回成功响应,不等待总结生成完成
|
||||
return new Result<Void>().ok(null);
|
||||
} catch (Exception e) {
|
||||
return new Result<Void>().error("启动异步总结生成任务失败: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
@PutMapping("/{id}")
|
||||
@Operation(summary = "更新智能体")
|
||||
@RequiresPermissions("sys:role:normal")
|
||||
@@ -186,6 +209,7 @@ public class AgentController {
|
||||
List<AgentChatHistoryDTO> result = agentChatHistoryService.getChatHistoryBySessionId(id, sessionId);
|
||||
return new Result<List<AgentChatHistoryDTO>>().ok(result);
|
||||
}
|
||||
|
||||
@GetMapping("/{id}/chat-history/user")
|
||||
@Operation(summary = "获取智能体聊天记录(用户)")
|
||||
@RequiresPermissions("sys:role:normal")
|
||||
|
||||
@@ -1,6 +1,9 @@
|
||||
package xiaozhi.modules.agent.dao;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import org.apache.ibatis.annotations.Mapper;
|
||||
import org.apache.ibatis.annotations.Param;
|
||||
|
||||
import com.baomidou.mybatisplus.core.mapper.BaseMapper;
|
||||
|
||||
@@ -15,12 +18,6 @@ import xiaozhi.modules.agent.entity.AgentChatHistoryEntity;
|
||||
*/
|
||||
@Mapper
|
||||
public interface AiAgentChatHistoryDao extends BaseMapper<AgentChatHistoryEntity> {
|
||||
/**
|
||||
* 根据智能体ID删除音频
|
||||
*
|
||||
* @param agentId 智能体ID
|
||||
*/
|
||||
void deleteAudioByAgentId(String agentId);
|
||||
|
||||
/**
|
||||
* 根据智能体ID删除聊天历史记录
|
||||
@@ -35,4 +32,19 @@ public interface AiAgentChatHistoryDao extends BaseMapper<AgentChatHistoryEntity
|
||||
* @param agentId 智能体ID
|
||||
*/
|
||||
void deleteAudioIdByAgentId(String agentId);
|
||||
|
||||
/**
|
||||
* 根据智能体ID获取所有音频ID列表
|
||||
*
|
||||
* @param agentId 智能体ID
|
||||
* @return 音频ID列表
|
||||
*/
|
||||
List<String> getAudioIdsByAgentId(String agentId);
|
||||
|
||||
/**
|
||||
* 批量删除音频
|
||||
*
|
||||
* @param audioIds 音频ID列表
|
||||
*/
|
||||
void deleteAudioByIds(@Param("audioIds") List<String> audioIds);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
package xiaozhi.modules.agent.dto;
|
||||
|
||||
import io.swagger.v3.oas.annotations.media.Schema;
|
||||
import lombok.Data;
|
||||
|
||||
/**
|
||||
* 智能体聊天记录总结DTO
|
||||
*/
|
||||
@Data
|
||||
@Schema(description = "智能体聊天记录总结对象")
|
||||
public class AgentChatSummaryDTO {
|
||||
|
||||
@Schema(description = "会话ID")
|
||||
private String sessionId;
|
||||
|
||||
@Schema(description = "智能体ID")
|
||||
private String agentId;
|
||||
|
||||
@Schema(description = "总结内容")
|
||||
private String summary;
|
||||
|
||||
@Schema(description = "总结状态")
|
||||
private boolean success;
|
||||
|
||||
@Schema(description = "错误信息")
|
||||
private String errorMessage;
|
||||
|
||||
public AgentChatSummaryDTO() {
|
||||
this.success = true;
|
||||
}
|
||||
|
||||
public AgentChatSummaryDTO(String sessionId, String agentId, String summary) {
|
||||
this.sessionId = sessionId;
|
||||
this.agentId = agentId;
|
||||
this.summary = summary;
|
||||
this.success = true;
|
||||
}
|
||||
|
||||
public AgentChatSummaryDTO(String sessionId, String errorMessage) {
|
||||
this.sessionId = sessionId;
|
||||
this.errorMessage = errorMessage;
|
||||
this.success = false;
|
||||
}
|
||||
|
||||
}
|
||||
@@ -0,0 +1,15 @@
|
||||
package xiaozhi.modules.agent.service;
|
||||
|
||||
/**
|
||||
* 智能体聊天记录总结服务接口
|
||||
*/
|
||||
public interface AgentChatSummaryService {
|
||||
|
||||
/**
|
||||
* 根据会话ID生成聊天记录总结并保存到智能体记忆
|
||||
*
|
||||
* @param sessionId 会话ID
|
||||
* @return 保存结果
|
||||
*/
|
||||
boolean generateAndSaveChatSummary(String sessionId);
|
||||
}
|
||||
@@ -17,6 +17,7 @@ import xiaozhi.modules.agent.entity.AgentChatHistoryEntity;
|
||||
import xiaozhi.modules.agent.entity.AgentEntity;
|
||||
import xiaozhi.modules.agent.service.AgentChatAudioService;
|
||||
import xiaozhi.modules.agent.service.AgentChatHistoryService;
|
||||
import xiaozhi.modules.agent.service.AgentChatSummaryService;
|
||||
import xiaozhi.modules.agent.service.AgentService;
|
||||
import xiaozhi.modules.agent.service.biz.AgentChatHistoryBizService;
|
||||
import xiaozhi.modules.device.entity.DeviceEntity;
|
||||
@@ -36,6 +37,7 @@ public class AgentChatHistoryBizServiceImpl implements AgentChatHistoryBizServic
|
||||
private final AgentService agentService;
|
||||
private final AgentChatHistoryService agentChatHistoryService;
|
||||
private final AgentChatAudioService agentChatAudioService;
|
||||
private final AgentChatSummaryService agentChatSummaryService;
|
||||
private final RedisUtils redisUtils;
|
||||
private final DeviceService deviceService;
|
||||
|
||||
@@ -50,7 +52,8 @@ public class AgentChatHistoryBizServiceImpl implements AgentChatHistoryBizServic
|
||||
public Boolean report(AgentChatHistoryReportDTO report) {
|
||||
String macAddress = report.getMacAddress();
|
||||
Byte chatType = report.getChatType();
|
||||
Long reportTimeMillis = null != report.getReportTime() ? report.getReportTime() * 1000 : System.currentTimeMillis();
|
||||
Long reportTimeMillis = null != report.getReportTime() ? report.getReportTime() * 1000
|
||||
: System.currentTimeMillis();
|
||||
log.info("小智设备聊天上报请求: macAddress={}, type={} reportTime={}", macAddress, chatType, reportTimeMillis);
|
||||
|
||||
// 根据设备MAC地址查询对应的默认智能体,判断是否需要上报
|
||||
@@ -105,7 +108,8 @@ public class AgentChatHistoryBizServiceImpl implements AgentChatHistoryBizServic
|
||||
/**
|
||||
* 组装上报数据
|
||||
*/
|
||||
private void saveChatText(AgentChatHistoryReportDTO report, String agentId, String macAddress, String audioId, Long reportTime) {
|
||||
private void saveChatText(AgentChatHistoryReportDTO report, String agentId, String macAddress, String audioId,
|
||||
Long reportTime) {
|
||||
// 构建聊天记录实体
|
||||
AgentChatHistoryEntity entity = AgentChatHistoryEntity.builder()
|
||||
.macAddress(macAddress)
|
||||
|
||||
@@ -84,7 +84,16 @@ public class AgentChatHistoryServiceImpl extends ServiceImpl<AiAgentChatHistoryD
|
||||
@Transactional(rollbackFor = Exception.class)
|
||||
public void deleteByAgentId(String agentId, Boolean deleteAudio, Boolean deleteText) {
|
||||
if (deleteAudio) {
|
||||
baseMapper.deleteAudioByAgentId(agentId);
|
||||
// 分批删除音频,避免超时
|
||||
List<String> audioIds = baseMapper.getAudioIdsByAgentId(agentId);
|
||||
if (audioIds != null && !audioIds.isEmpty()) {
|
||||
int batchSize = 1000; // 每批删除1000条
|
||||
for (int i = 0; i < audioIds.size(); i += batchSize) {
|
||||
int end = Math.min(i + batchSize, audioIds.size());
|
||||
List<String> batch = audioIds.subList(i, end);
|
||||
baseMapper.deleteAudioByIds(batch);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (deleteAudio && !deleteText) {
|
||||
baseMapper.deleteAudioIdByAgentId(agentId);
|
||||
@@ -107,7 +116,7 @@ public class AgentChatHistoryServiceImpl extends ServiceImpl<AiAgentChatHistoryD
|
||||
// 添加此行,确保查询结果按照创建时间降序排列
|
||||
// 使用id的原因:数据形式,id越大的创建时间就越晚,所以使用id的结果和创建时间降序排列结果一样
|
||||
// id作为降序排列的优势,性能高,有主键索引,不用在排序的时候重新进行排除扫描比较
|
||||
.orderByDesc(AgentChatHistoryEntity::getId);
|
||||
.orderByDesc(AgentChatHistoryEntity::getId);
|
||||
|
||||
// 构建分页查询,查询前50页数据
|
||||
Page<AgentChatHistoryEntity> pageParam = new Page<>(0, 50);
|
||||
|
||||
@@ -0,0 +1,423 @@
|
||||
package xiaozhi.modules.agent.service.impl;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import org.apache.commons.lang3.StringUtils;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import com.baomidou.mybatisplus.core.conditions.query.QueryWrapper;
|
||||
|
||||
import lombok.RequiredArgsConstructor;
|
||||
import xiaozhi.modules.agent.dto.AgentChatHistoryDTO;
|
||||
import xiaozhi.modules.agent.dto.AgentChatSummaryDTO;
|
||||
import xiaozhi.modules.agent.dto.AgentMemoryDTO;
|
||||
import xiaozhi.modules.agent.dto.AgentUpdateDTO;
|
||||
import xiaozhi.modules.agent.entity.AgentChatHistoryEntity;
|
||||
import xiaozhi.modules.agent.service.AgentChatHistoryService;
|
||||
import xiaozhi.modules.agent.service.AgentChatSummaryService;
|
||||
import xiaozhi.modules.agent.service.AgentService;
|
||||
import xiaozhi.modules.agent.vo.AgentInfoVO;
|
||||
import xiaozhi.modules.device.entity.DeviceEntity;
|
||||
import xiaozhi.modules.device.service.DeviceService;
|
||||
import xiaozhi.modules.llm.service.LLMService;
|
||||
import xiaozhi.modules.model.entity.ModelConfigEntity;
|
||||
import xiaozhi.modules.model.service.ModelConfigService;
|
||||
|
||||
/**
|
||||
* 智能体聊天记录总结服务实现类
|
||||
* 实现Python端mem_local_short.py中的总结逻辑
|
||||
*/
|
||||
@Service
|
||||
@RequiredArgsConstructor
|
||||
public class AgentChatSummaryServiceImpl implements AgentChatSummaryService {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(AgentChatSummaryServiceImpl.class);
|
||||
|
||||
private final AgentChatHistoryService agentChatHistoryService;
|
||||
private final AgentService agentService;
|
||||
private final DeviceService deviceService;
|
||||
private final LLMService llmService;
|
||||
private final ModelConfigService modelConfigService;
|
||||
|
||||
// 总结规则常量
|
||||
private static final int MAX_SUMMARY_LENGTH = 1800; // 最大总结长度
|
||||
private static final Pattern JSON_PATTERN = Pattern.compile("\\{.*?\\}", Pattern.DOTALL);
|
||||
private static final Pattern DEVICE_CONTROL_PATTERN = Pattern.compile("设备控制|设备操作|控制设备|设备状态",
|
||||
Pattern.CASE_INSENSITIVE);
|
||||
private static final Pattern WEATHER_PATTERN = Pattern.compile("天气|温度|湿度|降雨|气象", Pattern.CASE_INSENSITIVE);
|
||||
private static final Pattern DATE_PATTERN = Pattern.compile("日期|时间|星期|月份|年份", Pattern.CASE_INSENSITIVE);
|
||||
|
||||
private AgentChatSummaryDTO generateChatSummary(String sessionId) {
|
||||
try {
|
||||
System.out.println("开始生成会话 " + sessionId + " 的聊天记录总结");
|
||||
|
||||
// 1. 根据sessionId获取聊天记录
|
||||
List<AgentChatHistoryDTO> chatHistory = getChatHistoryBySessionId(sessionId);
|
||||
if (chatHistory == null || chatHistory.isEmpty()) {
|
||||
return new AgentChatSummaryDTO(sessionId, "未找到该会话的聊天记录");
|
||||
}
|
||||
|
||||
// 2. 获取智能体信息
|
||||
String agentId = getAgentIdFromSession(sessionId, chatHistory);
|
||||
if (StringUtils.isBlank(agentId)) {
|
||||
return new AgentChatSummaryDTO(sessionId, "无法获取智能体信息");
|
||||
}
|
||||
|
||||
// 3. 提取关键对话内容
|
||||
List<String> meaningfulMessages = extractMeaningfulMessages(chatHistory);
|
||||
if (meaningfulMessages.isEmpty()) {
|
||||
return new AgentChatSummaryDTO(sessionId, "没有有效的对话内容可总结");
|
||||
}
|
||||
|
||||
// 4. 生成总结(generateSummaryFromMessages方法已包含长度限制逻辑)
|
||||
String summary = generateSummaryFromMessages(meaningfulMessages, agentId);
|
||||
|
||||
System.out.println("成功生成会话 " + sessionId + " 的聊天记录总结,长度: " + summary.length() + " 字符");
|
||||
return new AgentChatSummaryDTO(sessionId, agentId, summary);
|
||||
|
||||
} catch (Exception e) {
|
||||
System.err.println("生成会话 " + sessionId + " 的聊天记录总结时发生错误: " + e.getMessage());
|
||||
return new AgentChatSummaryDTO(sessionId, "生成总结时发生错误: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean generateAndSaveChatSummary(String sessionId) {
|
||||
try {
|
||||
// 1. 生成总结
|
||||
AgentChatSummaryDTO summaryDTO = generateChatSummary(sessionId);
|
||||
if (!summaryDTO.isSuccess()) {
|
||||
System.err.println("生成总结失败: " + summaryDTO.getErrorMessage());
|
||||
return false;
|
||||
}
|
||||
|
||||
// 2. 获取设备信息(通过会话关联的设备)
|
||||
DeviceEntity device = getDeviceBySessionId(sessionId);
|
||||
if (device == null) {
|
||||
System.err.println("未找到与会话 " + sessionId + " 关联的设备");
|
||||
return false;
|
||||
}
|
||||
|
||||
// 3. 更新智能体记忆
|
||||
AgentMemoryDTO memoryDTO = new AgentMemoryDTO();
|
||||
memoryDTO.setSummaryMemory(summaryDTO.getSummary());
|
||||
|
||||
// 调用现有接口更新记忆
|
||||
agentService.updateAgentById(device.getAgentId(),
|
||||
new AgentUpdateDTO() {
|
||||
{
|
||||
setSummaryMemory(summaryDTO.getSummary());
|
||||
}
|
||||
});
|
||||
|
||||
System.out.println("成功保存会话 " + sessionId + " 的聊天记录总结到智能体 " + device.getAgentId());
|
||||
return true;
|
||||
|
||||
} catch (Exception e) {
|
||||
System.err.println("保存会话 " + sessionId + " 的聊天记录总结时发生错误: " + e.getMessage());
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 根据会话ID获取聊天记录
|
||||
*/
|
||||
private List<AgentChatHistoryDTO> getChatHistoryBySessionId(String sessionId) {
|
||||
try {
|
||||
// 这里需要根据sessionId获取聊天记录
|
||||
// 由于现有接口需要agentId,我们需要先找到关联的agentId
|
||||
String agentId = findAgentIdBySessionId(sessionId);
|
||||
if (StringUtils.isBlank(agentId)) {
|
||||
return null;
|
||||
}
|
||||
return agentChatHistoryService.getChatHistoryBySessionId(agentId, sessionId);
|
||||
} catch (Exception e) {
|
||||
System.err.println("获取会话 " + sessionId + " 的聊天记录失败: " + e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 根据会话ID查找关联的智能体ID
|
||||
*/
|
||||
private String findAgentIdBySessionId(String sessionId) {
|
||||
try {
|
||||
// 查询该会话的第一条记录获取agentId
|
||||
QueryWrapper<AgentChatHistoryEntity> wrapper = new QueryWrapper<>();
|
||||
wrapper.select("agent_id")
|
||||
.eq("session_id", sessionId)
|
||||
.last("LIMIT 1");
|
||||
|
||||
AgentChatHistoryEntity entity = agentChatHistoryService.getOne(wrapper);
|
||||
return entity != null ? entity.getAgentId() : null;
|
||||
} catch (Exception e) {
|
||||
System.err.println("根据会话ID " + sessionId + " 查找智能体ID失败: " + e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 从会话中获取智能体ID
|
||||
*/
|
||||
private String getAgentIdFromSession(String sessionId, List<AgentChatHistoryDTO> chatHistory) {
|
||||
// 直接从数据库查询智能体ID
|
||||
return findAgentIdBySessionId(sessionId);
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取有意义的对话内容(只提取用户消息,排除AI回复)
|
||||
*/
|
||||
private List<String> extractMeaningfulMessages(List<AgentChatHistoryDTO> chatHistory) {
|
||||
List<String> meaningfulMessages = new ArrayList<>();
|
||||
|
||||
for (AgentChatHistoryDTO message : chatHistory) {
|
||||
// 只处理用户消息(chatType = 1)
|
||||
if (message.getChatType() != null && message.getChatType() == 1) {
|
||||
String content = extractContentFromMessage(message);
|
||||
if (isMeaningfulMessage(content)) {
|
||||
meaningfulMessages.add(content);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return meaningfulMessages;
|
||||
}
|
||||
|
||||
/**
|
||||
* 从消息中提取内容(处理JSON格式)
|
||||
*/
|
||||
private String extractContentFromMessage(AgentChatHistoryDTO message) {
|
||||
String content = message.getContent();
|
||||
if (StringUtils.isBlank(content)) {
|
||||
return "";
|
||||
}
|
||||
|
||||
// 处理JSON格式内容(与前端ChatHistoryDialog.vue逻辑一致)
|
||||
Matcher matcher = JSON_PATTERN.matcher(content);
|
||||
if (matcher.find()) {
|
||||
String jsonContent = matcher.group();
|
||||
// 简化处理:提取JSON中的文本内容
|
||||
return extractTextFromJson(jsonContent);
|
||||
}
|
||||
|
||||
return content;
|
||||
}
|
||||
|
||||
/**
|
||||
* 从JSON中提取文本内容
|
||||
*/
|
||||
private String extractTextFromJson(String jsonContent) {
|
||||
// 简化处理:提取"content"字段的值
|
||||
Pattern contentPattern = Pattern.compile("\"content\"\s*:\s*\"([^\"]*)\"");
|
||||
Matcher matcher = contentPattern.matcher(jsonContent);
|
||||
if (matcher.find()) {
|
||||
return matcher.group(1);
|
||||
}
|
||||
return jsonContent;
|
||||
}
|
||||
|
||||
/**
|
||||
* 判断是否为有意义的消息
|
||||
*/
|
||||
private boolean isMeaningfulMessage(String content) {
|
||||
if (StringUtils.isBlank(content)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
// 排除设备控制信息
|
||||
if (DEVICE_CONTROL_PATTERN.matcher(content).find()) {
|
||||
return false;
|
||||
}
|
||||
|
||||
// 排除日期天气等无关内容
|
||||
if (WEATHER_PATTERN.matcher(content).find() || DATE_PATTERN.matcher(content).find()) {
|
||||
return false;
|
||||
}
|
||||
|
||||
// 排除过短的消息
|
||||
return content.length() >= 5;
|
||||
}
|
||||
|
||||
/**
|
||||
* 从消息生成总结
|
||||
*/
|
||||
private String generateSummaryFromMessages(List<String> messages, String agentId) {
|
||||
if (messages.isEmpty()) {
|
||||
return "本次对话内容较少,没有需要总结的重要信息。";
|
||||
}
|
||||
|
||||
// 构建完整的对话内容
|
||||
StringBuilder conversation = new StringBuilder();
|
||||
for (int i = 0; i < messages.size(); i++) {
|
||||
conversation.append("消息").append(i + 1).append(": ").append(messages.get(i)).append("\n");
|
||||
}
|
||||
|
||||
try {
|
||||
// 获取当前智能体的历史记忆
|
||||
String historyMemory = getCurrentAgentMemory(agentId);
|
||||
|
||||
// 调用LLM服务进行智能总结,传递agentId以获取正确的模型配置
|
||||
String summary = callJavaLLMForSummaryWithHistory(conversation.toString(), historyMemory, agentId);
|
||||
|
||||
// 应用总结规则:限制最大长度
|
||||
if (summary.length() > MAX_SUMMARY_LENGTH) {
|
||||
summary = summary.substring(0, MAX_SUMMARY_LENGTH) + "...";
|
||||
}
|
||||
|
||||
return summary;
|
||||
} catch (Exception e) {
|
||||
System.err.println("调用Java端LLM服务失败: " + e.getMessage());
|
||||
throw new RuntimeException("LLM服务不可用,无法生成聊天总结");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取当前智能体的历史记忆
|
||||
*/
|
||||
private String getCurrentAgentMemory(String agentId) {
|
||||
try {
|
||||
if (StringUtils.isBlank(agentId)) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 获取智能体信息
|
||||
AgentInfoVO agentInfo = agentService.getAgentById(agentId);
|
||||
if (agentInfo == null) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 返回智能体的当前总结记忆
|
||||
return agentInfo.getSummaryMemory();
|
||||
} catch (Exception e) {
|
||||
System.err.println("获取智能体历史记忆失败,agentId: " + agentId + ", 错误: " + e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 调用Java端LLM服务进行智能总结(支持历史记忆合并)
|
||||
*/
|
||||
private String callJavaLLMForSummaryWithHistory(String conversation, String historyMemory, String agentId) {
|
||||
try {
|
||||
// 获取智能体配置,从中提取记忆总结的模型ID
|
||||
String modelId = getMemorySummaryModelId(agentId);
|
||||
|
||||
if (StringUtils.isBlank(modelId)) {
|
||||
System.out.println("未找到记忆总结的LLM模型配置,使用默认LLM服务");
|
||||
return llmService.generateSummaryWithHistory(conversation, historyMemory, null, null);
|
||||
}
|
||||
|
||||
// 使用指定的模型ID调用LLM服务(支持历史记忆合并)
|
||||
String summary = llmService.generateSummaryWithHistory(conversation, historyMemory, null, modelId);
|
||||
|
||||
if (StringUtils.isNotBlank(summary) && !summary.equals("服务暂不可用") && !summary.equals("总结生成失败")) {
|
||||
return summary;
|
||||
}
|
||||
|
||||
throw new RuntimeException("Java端LLM服务返回异常: " + summary);
|
||||
|
||||
} catch (Exception e) {
|
||||
System.err.println("调用Java端LLM服务异常,agentId: " + agentId + ", 错误: " + e.getMessage());
|
||||
throw e;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 调用Java端LLM服务进行智能总结
|
||||
*/
|
||||
private String callJavaLLMForSummary(String conversation, String agentId) {
|
||||
try {
|
||||
// 获取智能体配置,从中提取记忆总结的模型ID
|
||||
String modelId = getMemorySummaryModelId(agentId);
|
||||
|
||||
if (StringUtils.isBlank(modelId)) {
|
||||
System.out.println("未找到记忆总结的LLM模型配置,使用默认LLM服务");
|
||||
return llmService.generateSummary(conversation);
|
||||
}
|
||||
|
||||
// 使用指定的模型ID调用LLM服务
|
||||
String summary = llmService.generateSummaryWithModel(conversation, modelId);
|
||||
|
||||
if (StringUtils.isNotBlank(summary) && !summary.equals("服务暂不可用") && !summary.equals("总结生成失败")) {
|
||||
return summary;
|
||||
}
|
||||
|
||||
throw new RuntimeException("Java端LLM服务返回异常: " + summary);
|
||||
|
||||
} catch (Exception e) {
|
||||
System.err.println("调用Java端LLM服务异常,agentId: " + agentId + ", 错误: " + e.getMessage());
|
||||
throw e;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取记忆总结的LLM模型ID
|
||||
*/
|
||||
private String getMemorySummaryModelId(String agentId) {
|
||||
try {
|
||||
if (StringUtils.isBlank(agentId)) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 获取智能体信息
|
||||
AgentInfoVO agentInfo = agentService.getAgentById(agentId);
|
||||
if (agentInfo == null) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 获取智能体的记忆模型ID
|
||||
String memModelId = agentInfo.getMemModelId();
|
||||
if (StringUtils.isBlank(memModelId)) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 获取记忆模型配置
|
||||
ModelConfigEntity memModelConfig = modelConfigService.getModelByIdFromCache(memModelId);
|
||||
if (memModelConfig == null || memModelConfig.getConfigJson() == null) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 从记忆模型配置中提取对应的LLM模型ID
|
||||
Map<String, Object> configMap = memModelConfig.getConfigJson();
|
||||
String llmModelId = (String) configMap.get("llm");
|
||||
|
||||
if (StringUtils.isBlank(llmModelId)) {
|
||||
// 如果记忆模型没有配置独立的LLM,则使用智能体的默认LLM模型
|
||||
return agentInfo.getLlmModelId();
|
||||
}
|
||||
|
||||
return llmModelId;
|
||||
} catch (Exception e) {
|
||||
System.err.println("获取记忆总结LLM模型ID失败,agentId: " + agentId + ", 错误: " + e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 根据会话ID获取设备信息
|
||||
*/
|
||||
private DeviceEntity getDeviceBySessionId(String sessionId) {
|
||||
try {
|
||||
// 查询该会话的第一条记录获取macAddress
|
||||
QueryWrapper<AgentChatHistoryEntity> wrapper = new QueryWrapper<>();
|
||||
wrapper.select("mac_address")
|
||||
.eq("session_id", sessionId)
|
||||
.last("LIMIT 1");
|
||||
|
||||
AgentChatHistoryEntity entity = agentChatHistoryService.getOne(wrapper);
|
||||
if (entity != null && StringUtils.isNotBlank(entity.getMacAddress())) {
|
||||
return deviceService.getDeviceByMacAddress(entity.getMacAddress());
|
||||
}
|
||||
return null;
|
||||
} catch (Exception e) {
|
||||
System.err.println("根据会话ID " + sessionId + " 查找设备信息失败: " + e.getMessage());
|
||||
return null;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -106,6 +106,15 @@ public class ConfigServiceImpl implements ConfigService {
|
||||
|
||||
@Override
|
||||
public Map<String, Object> getAgentModels(String macAddress, Map<String, String> selectedModule) {
|
||||
// 检查是否为管理控制台请求
|
||||
String redisKey = RedisKeys.getTmpRegisterMacKey(macAddress);
|
||||
Object isAdminRequest = redisUtils.get(redisKey);
|
||||
|
||||
if (isAdminRequest != null && "true".equals(isAdminRequest)) {
|
||||
// 管理控制台请求,返回getConfig的结果
|
||||
redisUtils.delete(redisKey); // 使用后清理
|
||||
return (Map<String, Object>) getConfig(true);
|
||||
}
|
||||
// 根据MAC地址查找设备
|
||||
DeviceEntity device = deviceService.getDeviceByMacAddress(macAddress);
|
||||
if (device == null) {
|
||||
|
||||
@@ -98,4 +98,14 @@ public interface DeviceService extends BaseService<DeviceEntity> {
|
||||
*/
|
||||
void updateDeviceConnectionInfo(String agentId, String deviceId, String appVersion);
|
||||
|
||||
/**
|
||||
* 生成WebSocket认证token
|
||||
*
|
||||
* @param clientId 客户端ID
|
||||
* @param username 用户名(通常为deviceId)
|
||||
* @return 认证token字符串
|
||||
* @throws Exception 生成token时的异常
|
||||
*/
|
||||
String generateWebSocketToken(String clientId, String username) throws Exception;
|
||||
|
||||
}
|
||||
@@ -518,7 +518,7 @@ public class DeviceServiceImpl extends BaseServiceImpl<DeviceDao, DeviceEntity>
|
||||
* @param username 用户名 (通常为deviceId/macAddress)
|
||||
* @return 认证token字符串
|
||||
*/
|
||||
private String generateWebSocketToken(String clientId, String username)
|
||||
public String generateWebSocketToken(String clientId, String username)
|
||||
throws NoSuchAlgorithmException, InvalidKeyException {
|
||||
// 从系统参数获取密钥
|
||||
String secretKey = sysParamsService.getValue(Constant.SERVER_SECRET, false);
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
package xiaozhi.modules.llm.service;
|
||||
|
||||
/**
|
||||
* LLM服务接口
|
||||
* 支持多种大模型调用
|
||||
*/
|
||||
public interface LLMService {
|
||||
|
||||
/**
|
||||
* 生成聊天记录总结
|
||||
*
|
||||
* @param conversation 对话内容
|
||||
* @param promptTemplate 提示词模板
|
||||
* @return 总结结果
|
||||
*/
|
||||
String generateSummary(String conversation, String promptTemplate);
|
||||
|
||||
/**
|
||||
* 生成聊天记录总结(使用默认提示词)
|
||||
*
|
||||
* @param conversation 对话内容
|
||||
* @return 总结结果
|
||||
*/
|
||||
String generateSummary(String conversation);
|
||||
|
||||
/**
|
||||
* 生成聊天记录总结(指定模型ID)
|
||||
*
|
||||
* @param conversation 对话内容
|
||||
* @param modelId 模型ID
|
||||
* @return 总结结果
|
||||
*/
|
||||
String generateSummaryWithModel(String conversation, String modelId);
|
||||
|
||||
/**
|
||||
* 生成聊天记录总结(指定模型ID和提示词模板)
|
||||
*
|
||||
* @param conversation 对话内容
|
||||
* @param promptTemplate 提示词模板
|
||||
* @param modelId 模型ID
|
||||
* @return 总结结果
|
||||
*/
|
||||
String generateSummary(String conversation, String promptTemplate, String modelId);
|
||||
|
||||
/**
|
||||
* 生成聊天记录总结(包含历史记忆合并)
|
||||
*
|
||||
* @param conversation 对话内容
|
||||
* @param historyMemory 历史记忆
|
||||
* @param promptTemplate 提示词模板
|
||||
* @param modelId 模型ID
|
||||
* @return 总结结果
|
||||
*/
|
||||
String generateSummaryWithHistory(String conversation, String historyMemory, String promptTemplate, String modelId);
|
||||
|
||||
/**
|
||||
* 检查服务是否可用
|
||||
*
|
||||
* @return 是否可用
|
||||
*/
|
||||
boolean isAvailable();
|
||||
|
||||
/**
|
||||
* 检查指定模型的服务是否可用
|
||||
*
|
||||
* @param modelId 模型ID
|
||||
* @return 是否可用
|
||||
*/
|
||||
boolean isAvailable(String modelId);
|
||||
}
|
||||
@@ -0,0 +1,305 @@
|
||||
package xiaozhi.modules.llm.service.impl;
|
||||
|
||||
import java.util.HashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import org.apache.commons.lang3.StringUtils;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.http.HttpEntity;
|
||||
import org.springframework.http.HttpHeaders;
|
||||
import org.springframework.http.HttpMethod;
|
||||
import org.springframework.http.MediaType;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.stereotype.Service;
|
||||
import org.springframework.web.client.RestTemplate;
|
||||
|
||||
import cn.hutool.json.JSONArray;
|
||||
import cn.hutool.json.JSONObject;
|
||||
import cn.hutool.json.JSONUtil;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import xiaozhi.modules.llm.service.LLMService;
|
||||
import xiaozhi.modules.model.entity.ModelConfigEntity;
|
||||
import xiaozhi.modules.model.service.ModelConfigService;
|
||||
|
||||
/**
|
||||
* OpenAI风格API的LLM服务实现
|
||||
* 支持阿里云、DeepSeek、ChatGLM等兼容OpenAI API的模型
|
||||
*/
|
||||
@Slf4j
|
||||
@Service
|
||||
public class OpenAIStyleLLMServiceImpl implements LLMService {
|
||||
|
||||
@Autowired
|
||||
private ModelConfigService modelConfigService;
|
||||
|
||||
private final RestTemplate restTemplate = new RestTemplate();
|
||||
|
||||
private static final String DEFAULT_SUMMARY_PROMPT = "你是一个经验丰富的记忆总结者,擅长将对话内容进行总结摘要,遵循以下规则:\n1、总结用户的重要信息,以便在未来的对话中提供更个性化的服务\n2、不要重复总结,不要遗忘之前记忆,除非原来的记忆超过了1800字,否则不要遗忘、不要压缩用户的历史记忆\n3、用户操控的设备音量、播放音乐、天气、退出、不想对话等和用户本身无关的内容,这些信息不需要加入到总结中\n4、聊天内容中的今天的日期时间、今天的天气情况与用户事件无关的数据,这些信息如果当成记忆存储会影响后续对话,这些信息不需要加入到总结中\n5、不要把设备操控的成果结果和失败结果加入到总结中,也不要把用户的一些废话加入到总结中\n6、不要为了总结而总结,如果用户的聊天没有意义,请返回原来的历史记录也是可以的\n7、只需要返回总结摘要,严格控制在1800字内\n8、不要包含代码、xml,不需要解释、注释和说明,保存记忆时仅从对话提取信息,不要混入示例内容\n9、如果提供了历史记忆,请将新对话内容与历史记忆进行智能合并,保留有价值的历史信息,同时添加新的重要信息\n\n历史记忆:\n{history_memory}\n\n新对话内容:\n{conversation}";
|
||||
|
||||
@Override
|
||||
public String generateSummary(String conversation) {
|
||||
return generateSummary(conversation, null, null);
|
||||
}
|
||||
|
||||
@Override
|
||||
public String generateSummaryWithModel(String conversation, String modelId) {
|
||||
return generateSummary(conversation, null, modelId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public String generateSummary(String conversation, String promptTemplate, String modelId) {
|
||||
if (!isAvailable()) {
|
||||
log.warn("LLM服务不可用,无法生成总结");
|
||||
return "LLM服务不可用,无法生成总结";
|
||||
}
|
||||
|
||||
try {
|
||||
// 从智控台获取LLM模型配置
|
||||
ModelConfigEntity llmConfig;
|
||||
if (modelId != null && !modelId.trim().isEmpty()) {
|
||||
// 通过具体模型ID获取配置
|
||||
llmConfig = modelConfigService.getModelByIdFromCache(modelId);
|
||||
} else {
|
||||
// 保持向后兼容,使用默认配置
|
||||
llmConfig = getDefaultLLMConfig();
|
||||
}
|
||||
|
||||
if (llmConfig == null || llmConfig.getConfigJson() == null) {
|
||||
log.error("未找到可用的LLM模型配置,modelId: {}", modelId);
|
||||
return "未找到可用的LLM模型配置";
|
||||
}
|
||||
|
||||
JSONObject configJson = llmConfig.getConfigJson();
|
||||
String baseUrl = configJson.getStr("base_url");
|
||||
String model = configJson.getStr("model_name");
|
||||
String apiKey = configJson.getStr("api_key");
|
||||
Double temperature = configJson.getDouble("temperature");
|
||||
Integer maxTokens = configJson.getInt("max_tokens");
|
||||
|
||||
if (StringUtils.isBlank(baseUrl) || StringUtils.isBlank(apiKey)) {
|
||||
log.error("LLM配置不完整,baseUrl或apiKey为空");
|
||||
return "LLM配置不完整,无法生成总结";
|
||||
}
|
||||
|
||||
// 构建提示词
|
||||
String prompt = (promptTemplate != null ? promptTemplate : DEFAULT_SUMMARY_PROMPT).replace("{conversation}",
|
||||
conversation);
|
||||
|
||||
// 构建请求体
|
||||
Map<String, Object> requestBody = new HashMap<>();
|
||||
requestBody.put("model", model != null ? model : "gpt-3.5-turbo");
|
||||
|
||||
Map<String, Object>[] messages = new Map[1];
|
||||
Map<String, Object> message = new HashMap<>();
|
||||
message.put("role", "user");
|
||||
message.put("content", prompt);
|
||||
messages[0] = message;
|
||||
|
||||
requestBody.put("messages", messages);
|
||||
requestBody.put("temperature", temperature != null ? temperature : 0.7);
|
||||
requestBody.put("max_tokens", maxTokens != null ? maxTokens : 2000);
|
||||
|
||||
// 发送HTTP请求
|
||||
HttpHeaders headers = new HttpHeaders();
|
||||
headers.setContentType(MediaType.APPLICATION_JSON);
|
||||
headers.set("Authorization", "Bearer " + apiKey);
|
||||
|
||||
HttpEntity<Map<String, Object>> entity = new HttpEntity<>(requestBody, headers);
|
||||
|
||||
// 构建完整的API URL
|
||||
String apiUrl = baseUrl;
|
||||
if (!apiUrl.endsWith("/chat/completions")) {
|
||||
if (!apiUrl.endsWith("/")) {
|
||||
apiUrl += "/";
|
||||
}
|
||||
apiUrl += "chat/completions";
|
||||
}
|
||||
|
||||
ResponseEntity<String> response = restTemplate.exchange(
|
||||
apiUrl, HttpMethod.POST, entity, String.class);
|
||||
|
||||
if (response.getStatusCode().is2xxSuccessful()) {
|
||||
JSONObject responseJson = JSONUtil.parseObj(response.getBody());
|
||||
JSONArray choices = responseJson.getJSONArray("choices");
|
||||
if (choices != null && choices.size() > 0) {
|
||||
JSONObject choice = choices.getJSONObject(0);
|
||||
JSONObject messageObj = choice.getJSONObject("message");
|
||||
return messageObj.getStr("content");
|
||||
}
|
||||
} else {
|
||||
log.error("LLM API调用失败,状态码:{},响应:{}", response.getStatusCode(), response.getBody());
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.error("调用LLM服务生成总结时发生异常,modelId: {}", modelId, e);
|
||||
}
|
||||
|
||||
return "生成总结失败,请稍后重试";
|
||||
}
|
||||
|
||||
@Override
|
||||
public String generateSummary(String conversation, String promptTemplate) {
|
||||
return generateSummary(conversation, promptTemplate, null);
|
||||
}
|
||||
|
||||
@Override
|
||||
public String generateSummaryWithHistory(String conversation, String historyMemory, String promptTemplate,
|
||||
String modelId) {
|
||||
if (!isAvailable()) {
|
||||
log.warn("LLM服务不可用,无法生成总结");
|
||||
return "LLM服务不可用,无法生成总结";
|
||||
}
|
||||
|
||||
try {
|
||||
// 从智控台获取LLM模型配置
|
||||
ModelConfigEntity llmConfig;
|
||||
if (modelId != null && !modelId.trim().isEmpty()) {
|
||||
// 通过具体模型ID获取配置
|
||||
llmConfig = modelConfigService.getModelByIdFromCache(modelId);
|
||||
} else {
|
||||
// 保持向后兼容,使用默认配置
|
||||
llmConfig = getDefaultLLMConfig();
|
||||
}
|
||||
|
||||
if (llmConfig == null || llmConfig.getConfigJson() == null) {
|
||||
log.error("未找到可用的LLM模型配置,modelId: {}", modelId);
|
||||
return "未找到可用的LLM模型配置";
|
||||
}
|
||||
|
||||
JSONObject configJson = llmConfig.getConfigJson();
|
||||
String baseUrl = configJson.getStr("base_url");
|
||||
String model = configJson.getStr("model_name");
|
||||
String apiKey = configJson.getStr("api_key");
|
||||
|
||||
if (StringUtils.isBlank(baseUrl) || StringUtils.isBlank(apiKey)) {
|
||||
log.error("LLM配置不完整,baseUrl或apiKey为空");
|
||||
return "LLM配置不完整,无法生成总结";
|
||||
}
|
||||
|
||||
// 构建提示词,包含历史记忆
|
||||
String prompt = (promptTemplate != null ? promptTemplate : DEFAULT_SUMMARY_PROMPT)
|
||||
.replace("{history_memory}", historyMemory != null ? historyMemory : "无历史记忆")
|
||||
.replace("{conversation}", conversation);
|
||||
|
||||
// 构建请求体
|
||||
Map<String, Object> requestBody = new HashMap<>();
|
||||
requestBody.put("model", model != null ? model : "gpt-3.5-turbo");
|
||||
|
||||
Map<String, Object>[] messages = new Map[1];
|
||||
Map<String, Object> message = new HashMap<>();
|
||||
message.put("role", "user");
|
||||
message.put("content", prompt);
|
||||
messages[0] = message;
|
||||
|
||||
requestBody.put("messages", messages);
|
||||
requestBody.put("temperature", 0.2);
|
||||
requestBody.put("max_tokens", 2000);
|
||||
|
||||
// 发送HTTP请求
|
||||
HttpHeaders headers = new HttpHeaders();
|
||||
headers.setContentType(MediaType.APPLICATION_JSON);
|
||||
headers.set("Authorization", "Bearer " + apiKey);
|
||||
|
||||
HttpEntity<Map<String, Object>> entity = new HttpEntity<>(requestBody, headers);
|
||||
|
||||
// 构建完整的API URL
|
||||
String apiUrl = baseUrl;
|
||||
if (!apiUrl.endsWith("/chat/completions")) {
|
||||
if (!apiUrl.endsWith("/")) {
|
||||
apiUrl += "/";
|
||||
}
|
||||
apiUrl += "chat/completions";
|
||||
}
|
||||
|
||||
ResponseEntity<String> response = restTemplate.exchange(
|
||||
apiUrl, HttpMethod.POST, entity, String.class);
|
||||
|
||||
if (response.getStatusCode().is2xxSuccessful()) {
|
||||
JSONObject responseJson = JSONUtil.parseObj(response.getBody());
|
||||
JSONArray choices = responseJson.getJSONArray("choices");
|
||||
if (choices != null && choices.size() > 0) {
|
||||
JSONObject choice = choices.getJSONObject(0);
|
||||
JSONObject messageObj = choice.getJSONObject("message");
|
||||
return messageObj.getStr("content");
|
||||
}
|
||||
} else {
|
||||
log.error("LLM API调用失败,状态码:{},响应:{}", response.getStatusCode(), response.getBody());
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.error("调用LLM服务生成总结时发生异常,modelId: {}", modelId, e);
|
||||
}
|
||||
|
||||
return "生成总结失败,请稍后重试";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean isAvailable() {
|
||||
try {
|
||||
ModelConfigEntity defaultLLMConfig = getDefaultLLMConfig();
|
||||
if (defaultLLMConfig == null || defaultLLMConfig.getConfigJson() == null) {
|
||||
return false;
|
||||
}
|
||||
|
||||
JSONObject configJson = defaultLLMConfig.getConfigJson();
|
||||
String baseUrl = configJson.getStr("base_url");
|
||||
String apiKey = configJson.getStr("api_key");
|
||||
|
||||
return baseUrl != null && !baseUrl.trim().isEmpty() &&
|
||||
apiKey != null && !apiKey.trim().isEmpty();
|
||||
} catch (Exception e) {
|
||||
log.error("检查LLM服务可用性时发生异常:", e);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean isAvailable(String modelId) {
|
||||
try {
|
||||
if (modelId == null || modelId.trim().isEmpty()) {
|
||||
return isAvailable();
|
||||
}
|
||||
|
||||
// 通过具体模型ID获取配置
|
||||
ModelConfigEntity modelConfig = modelConfigService.getModelByIdFromCache(modelId);
|
||||
if (modelConfig == null || modelConfig.getConfigJson() == null) {
|
||||
log.warn("未找到指定的LLM模型配置,modelId: {}", modelId);
|
||||
return false;
|
||||
}
|
||||
|
||||
JSONObject configJson = modelConfig.getConfigJson();
|
||||
String baseUrl = configJson.getStr("base_url");
|
||||
String apiKey = configJson.getStr("api_key");
|
||||
|
||||
return baseUrl != null && !baseUrl.trim().isEmpty() &&
|
||||
apiKey != null && !apiKey.trim().isEmpty();
|
||||
} catch (Exception e) {
|
||||
log.error("检查LLM服务可用性时发生异常,modelId: {}", modelId, e);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 从智控台获取默认的LLM模型配置
|
||||
*/
|
||||
private ModelConfigEntity getDefaultLLMConfig() {
|
||||
try {
|
||||
// 获取所有启用的LLM模型配置
|
||||
List<ModelConfigEntity> llmConfigs = modelConfigService.getEnabledModelsByType("LLM");
|
||||
if (llmConfigs == null || llmConfigs.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// 优先返回默认配置,如果没有默认配置则返回第一个启用的配置
|
||||
for (ModelConfigEntity config : llmConfigs) {
|
||||
if (config.getIsDefault() != null && config.getIsDefault() == 1) {
|
||||
return config;
|
||||
}
|
||||
}
|
||||
|
||||
return llmConfigs.get(0);
|
||||
} catch (Exception e) {
|
||||
log.error("获取LLM模型配置时发生异常:", e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -55,4 +55,12 @@ public interface ModelConfigService extends BaseService<ModelConfigEntity> {
|
||||
* @return TTS平台列表(id和modelName)
|
||||
*/
|
||||
List<Map<String, Object>> getTtsPlatformList();
|
||||
|
||||
/**
|
||||
* 根据模型类型获取所有启用的模型配置
|
||||
*
|
||||
* @param modelType 模型类型(如:LLM, TTS, ASR等)
|
||||
* @return 启用的模型配置列表
|
||||
*/
|
||||
List<ModelConfigEntity> getEnabledModelsByType(String modelType);
|
||||
}
|
||||
|
||||
@@ -502,4 +502,22 @@ public class ModelConfigServiceImpl extends BaseServiceImpl<ModelConfigDao, Mode
|
||||
public List<Map<String, Object>> getTtsPlatformList() {
|
||||
return modelConfigDao.getTtsPlatformList();
|
||||
}
|
||||
|
||||
/**
|
||||
* 根据模型类型获取所有启用的模型配置
|
||||
*/
|
||||
@Override
|
||||
public List<ModelConfigEntity> getEnabledModelsByType(String modelType) {
|
||||
if (StringUtils.isBlank(modelType)) {
|
||||
return null;
|
||||
}
|
||||
|
||||
List<ModelConfigEntity> entities = modelConfigDao.selectList(
|
||||
new QueryWrapper<ModelConfigEntity>()
|
||||
.eq("model_type", modelType)
|
||||
.eq("is_enabled", 1)
|
||||
.orderByAsc("sort"));
|
||||
|
||||
return entities;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -89,7 +89,7 @@ public class ShiroConfig {
|
||||
filterMap.put("/config/**", "server");
|
||||
filterMap.put("/agent/chat-history/report", "server");
|
||||
filterMap.put("/agent/chat-history/download/**", "anon");
|
||||
filterMap.put("/agent/saveMemory/**", "server");
|
||||
filterMap.put("/agent/chat-summary/**", "server");
|
||||
filterMap.put("/agent/play/**", "anon");
|
||||
filterMap.put("/voiceClone/play/**", "anon");
|
||||
filterMap.put("/**", "oauth2");
|
||||
|
||||
@@ -31,6 +31,8 @@ import xiaozhi.modules.sys.dto.ServerActionResponseDTO;
|
||||
import xiaozhi.modules.sys.enums.ServerActionEnum;
|
||||
import xiaozhi.modules.sys.service.SysParamsService;
|
||||
import xiaozhi.modules.sys.utils.WebSocketClientManager;
|
||||
import xiaozhi.modules.device.service.DeviceService;
|
||||
import xiaozhi.common.redis.RedisUtils;
|
||||
|
||||
/**
|
||||
* 服务端管理控制器
|
||||
@@ -41,6 +43,8 @@ import xiaozhi.modules.sys.utils.WebSocketClientManager;
|
||||
@AllArgsConstructor
|
||||
public class ServerSideManageController {
|
||||
private final SysParamsService sysParamsService;
|
||||
private final DeviceService deviceService;
|
||||
private final RedisUtils redisUtils;
|
||||
private static final ObjectMapper objectMapper;
|
||||
static {
|
||||
objectMapper = new ObjectMapper();
|
||||
@@ -85,9 +89,22 @@ public class ServerSideManageController {
|
||||
return false;
|
||||
}
|
||||
String serverSK = sysParamsService.getValue(Constant.SERVER_SECRET, true);
|
||||
|
||||
String deviceId = UUID.randomUUID().toString();
|
||||
String clientId = UUID.randomUUID().toString();
|
||||
|
||||
String redisKey = xiaozhi.common.redis.RedisKeys.getTmpRegisterMacKey(deviceId);
|
||||
redisUtils.set(redisKey, "true", 300); // 5分钟有效期
|
||||
|
||||
WebSocketHttpHeaders headers = new WebSocketHttpHeaders();
|
||||
headers.add("device-id", UUID.randomUUID().toString());
|
||||
headers.add("client-id", UUID.randomUUID().toString());
|
||||
headers.add("device-id", deviceId);
|
||||
headers.add("client-id", clientId);
|
||||
try {
|
||||
String token = deviceService.generateWebSocketToken(clientId, deviceId);
|
||||
headers.add("authorization", "Bearer " + token);
|
||||
} catch (Exception e) {
|
||||
throw new RenException(ErrorCode.WEB_SOCKET_CONNECT_FAILED);
|
||||
}
|
||||
|
||||
try (WebSocketClientManager client = new WebSocketClientManager.Builder()
|
||||
.connectTimeout(3, TimeUnit.SECONDS)
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
-- 更新HuoshanDoubleStreamTTS供应器配置,增加开启链接复用选项
|
||||
UPDATE `ai_model_provider`
|
||||
SET fields = '[{"key": "ws_url", "type": "string", "label": "WebSocket地址"}, {"key": "appid", "type": "string", "label": "应用ID"}, {"key": "access_token", "type": "string", "label": "访问令牌"}, {"key": "resource_id", "type": "string", "label": "资源ID"}, {"key": "speaker", "type": "string", "label": "默认音色"}, {"key": "enable_ws_reuse", "type": "boolean", "label": "是否开启链接复用", "default": true}, {"key": "speech_rate", "type": "number", "label": "语速(-50~100)"}, {"key": "loudness_rate", "type": "number", "label": "音量(-50~100)"}, {"key": "pitch", "type": "number", "label": "音高(-12~12)"}]'
|
||||
WHERE id = 'SYSTEM_TTS_HSDSTTS';
|
||||
|
||||
UPDATE `ai_model_config` SET
|
||||
`doc_link` = 'https://console.volcengine.com/speech/service/10007',
|
||||
`remark` = '火山引擎语音合成服务配置说明:
|
||||
1. 访问 https://www.volcengine.com/ 注册并开通火山引擎账号
|
||||
2. 访问 https://console.volcengine.com/speech/service/10007 开通语音合成大模型,购买音色
|
||||
3. 在页面底部获取appid和access_token
|
||||
5. 资源ID固定为:volc.service_type.10029(大模型语音合成及混音)
|
||||
6. 链接复用:开启WebSocket连接复用,默认true减少链接损耗(注意:复用后设备处于聆听状态时空闲链接会占并发数)
|
||||
7. 语速:-50~100,可不填,正常默认值0,可填-50~100
|
||||
8. 音量:-50~100,可不填,正常默认值0,可填-50~100
|
||||
9. 音高:-12~12,可不填,正常默认值0,可填-12~12
|
||||
10. 填入配置文件中' WHERE `id` = 'TTS_HuoshanDoubleStreamTTS';
|
||||
@@ -0,0 +1 @@
|
||||
INSERT INTO `sys_params` (id, param_code, param_value, value_type, param_type, remark) VALUES (311, 'enable_websocket_ping', 'false', 'boolean', 1, '是否启用WebSocket心跳保活机制');
|
||||
@@ -0,0 +1,2 @@
|
||||
-- 为智能体聊天历史记录添加音频ID索引
|
||||
ALTER TABLE ai_agent_chat_history ADD INDEX idx_ai_agent_chat_history_audio_id (audio_id);
|
||||
@@ -0,0 +1,10 @@
|
||||
-- 更新豆包流式ASR供应器,增加end_window_size配置
|
||||
delete from `ai_model_provider` where id = 'SYSTEM_ASR_DoubaoStreamASR';
|
||||
INSERT INTO `ai_model_provider` (`id`, `model_type`, `provider_code`, `name`, `fields`, `sort`, `creator`, `create_date`, `updater`, `update_date`) VALUES
|
||||
('SYSTEM_ASR_DoubaoStreamASR', 'ASR', 'doubao_stream', '火山引擎语音识别(流式)', '[{"key":"appid","label":"应用ID","type":"string"},{"key":"access_token","label":"访问令牌","type":"string"},{"key":"cluster","label":"集群","type":"string"},{"key":"boosting_table_name","label":"热词文件名称","type":"string"},{"key":"correct_table_name","label":"替换词文件名称","type":"string"},{"key":"output_dir","label":"输出目录","type":"string"},{"key":"end_window_size","label":"静音判定时长(ms)","type":"number"}]', 3, 1, NOW(), 1, NOW());
|
||||
|
||||
|
||||
-- 更新豆包流式ASR模型配置,增加end_window_size默认值
|
||||
UPDATE `ai_model_config` SET
|
||||
`config_json` = JSON_SET(`config_json`, '$.end_window_size', 200)
|
||||
WHERE `id` = 'ASR_DoubaoStreamASR' AND JSON_EXTRACT(`config_json`, '$.end_window_size') IS NULL;
|
||||
@@ -0,0 +1,94 @@
|
||||
-- 添加阿里百炼Paraformer实时语音识别服务配置
|
||||
delete from `ai_model_provider` where id = 'SYSTEM_ASR_AliyunBLStream';
|
||||
INSERT INTO `ai_model_provider` (`id`, `model_type`, `provider_code`, `name`, `fields`, `sort`, `creator`, `create_date`, `updater`, `update_date`) VALUES
|
||||
('SYSTEM_ASR_AliyunBLStream', 'ASR', 'aliyunbl_stream', '阿里百炼Paraformer实时语音识别', '[{"key":"api_key","label":"API密钥","type":"password"},{"key":"model","label":"模型名称","type":"string"},{"key":"format","label":"音频格式","type":"string"},{"key":"sample_rate","label":"采样率","type":"number"},{"key":"output_dir","label":"输出目录","type":"string"}]', 18, 1, NOW(), 1, NOW());
|
||||
|
||||
delete from `ai_model_config` where id = 'ASR_AliyunBLStream';
|
||||
INSERT INTO `ai_model_config` VALUES ('ASR_AliyunBLStream', 'ASR', 'AliyunBLStream', '阿里百炼Paraformer实时语音识别', 0, 1, '{"type": "aliyunbl_stream", "api_key": "", "model": "paraformer-realtime-v2", "format": "pcm", "sample_rate": 16000, "disfluency_removal_enabled": false, "semantic_punctuation_enabled": false, "max_sentence_silence": 200, "multi_threshold_mode_enabled": false, "punctuation_prediction_enabled": true, "inverse_text_normalization_enabled": true, "output_dir": "tmp/"}', 'https://help.aliyun.com/zh/model-studio/websocket-for-paraformer-real-time-service', '支持多语言、热词定制、语义断句等高级功能', 21, NULL, NULL, NULL, NULL);
|
||||
|
||||
-- 更新阿里百炼Paraformer模型配置的说明文档
|
||||
UPDATE `ai_model_config` SET
|
||||
`doc_link` = 'https://help.aliyun.com/zh/model-studio/websocket-for-paraformer-real-time-service',
|
||||
`remark` = '阿里百炼Paraformer实时语音识别配置说明:
|
||||
1. 登录阿里云百炼平台 https://bailian.console.aliyun.com/
|
||||
2. 创建API-KEY https://bailian.console.aliyun.com/#/api-key
|
||||
3. 支持模型:paraformer-realtime-v2(推荐)、paraformer-realtime-8k-v2、paraformer-realtime-v1、paraformer-realtime-8k-v1
|
||||
4. 功能特性:
|
||||
- 多语言支持(中文含方言、英文、日语、韩语、德语、法语、俄语)
|
||||
- 热词定制(vocabulary_id参数),详细说明请参考:https://help.aliyun.com/zh/model-studio/custom-hot-words?
|
||||
- 语义断句/VAD断句(semantic_punctuation_enabled参数)
|
||||
- 自动标点符号、ITN、过滤语气词等
|
||||
5. 参数说明:
|
||||
- model: 模型名称,推荐paraformer-realtime-v2
|
||||
- sample_rate: 采样率(Hz),v2支持任意采样率,v1仅支持16000,8k版本仅支持8000
|
||||
- semantic_punctuation_enabled: false为VAD断句(低延迟),true为语义断句(高准确)
|
||||
- max_sentence_silence: VAD断句静音时长阈值(200-6000ms)
|
||||
' WHERE `id` = 'ASR_AliyunBLStream';
|
||||
|
||||
|
||||
-- 更新豆包流式ASR供应器,增加配置
|
||||
delete from `ai_model_provider` where id = 'SYSTEM_ASR_DoubaoStreamASR';
|
||||
INSERT INTO `ai_model_provider` (`id`, `model_type`, `provider_code`, `name`, `fields`, `sort`, `creator`, `create_date`, `updater`, `update_date`) VALUES
|
||||
('SYSTEM_ASR_DoubaoStreamASR', 'ASR', 'doubao_stream', '火山引擎语音识别(流式)', '[{"key":"appid","label":"应用ID","type":"string"},{"key":"access_token","label":"访问令牌","type":"string"},{"key":"cluster","label":"集群","type":"string"},{"key":"boosting_table_name","label":"热词文件名称","type":"string"},{"key":"correct_table_name","label":"替换词文件名称","type":"string"},{"key":"output_dir","label":"输出目录","type":"string"},{"key":"end_window_size","label":"静音判定时长(ms)","type":"number"},{"key":"enable_multilingual","label":"是否开启多语种识别模式","type":"boolean"},{"key":"language","label":"指定语言编码","type":"string"}]', 3, 1, NOW(), 1, NOW());
|
||||
UPDATE `ai_model_config` SET
|
||||
`remark` = '豆包ASR配置说明:
|
||||
1. 豆包ASR和豆包(流式)ASR的区别是:豆包ASR是按次收费,豆包(流式)ASR是按时收费
|
||||
2. 一般来说按次收费的更便宜,但是豆包(流式)ASR使用了大模型技术,效果更好
|
||||
3. 需要在火山引擎控制台创建应用并获取appid和access_token
|
||||
4. 支持中文语音识别
|
||||
5. 需要网络连接
|
||||
6. 输出文件保存在tmp/目录
|
||||
申请步骤:
|
||||
1. 访问 https://console.volcengine.com/speech/app
|
||||
2. 创建新应用
|
||||
3. 获取appid和access_token
|
||||
4. 填入配置文件中
|
||||
如需设置热词,请参考:https://www.volcengine.com/docs/6561/155738
|
||||
如开启多语种识别模式,请设置language当该键为空时,该模型支持中英文、上海话、闽南语,四川、陕西、粤语识别。其他语种请参考:https://www.volcengine.com/docs/6561/1354869
|
||||
' WHERE `id` = 'ASR_DoubaoStreamASR';
|
||||
|
||||
-- 更新豆包流式ASR模型配置,增加enable_multilingual默认值
|
||||
UPDATE `ai_model_config` SET
|
||||
`config_json` = JSON_SET(
|
||||
`config_json`,
|
||||
'$.enable_multilingual', false,
|
||||
'$.language', 'zh-CN'
|
||||
)
|
||||
WHERE `id` = 'ASR_DoubaoStreamASR'
|
||||
AND JSON_EXTRACT(`config_json`, '$.enable_multilingual') IS NULL
|
||||
AND JSON_EXTRACT(`config_json`, '$.language') IS NULL;
|
||||
|
||||
|
||||
-- 更新HuoshanDoubleStreamTTS供应器配置,增加多情感音色参数
|
||||
UPDATE `ai_model_provider`
|
||||
SET `fields` = '[{"key": "ws_url", "type": "string", "label": "WebSocket地址"}, {"key": "appid", "type": "string", "label": "应用ID"}, {"key": "access_token", "type": "string", "label": "访问令牌"}, {"key": "resource_id", "type": "string", "label": "资源ID"}, {"key": "speaker", "type": "string", "label": "默认音色"}, {"key": "enable_ws_reuse", "type": "boolean", "label": "是否开启链接复用", "default": true}, {"key": "speech_rate", "type": "number", "label": "语速(-50~100)"}, {"key": "loudness_rate", "type": "number", "label": "音量(-50~100)"}, {"key": "pitch", "type": "number", "label": "音高(-12~12)"}, {"key": "emotion_scale", "type": "number", "label": "情感强度(1-5)"}, {"key": "emotion", "type": "string", "label": "情感类型"}]'
|
||||
WHERE `id` = 'SYSTEM_TTS_HSDSTTS';
|
||||
|
||||
-- 更新默认值
|
||||
UPDATE `ai_model_config` SET
|
||||
`config_json` = JSON_SET(
|
||||
`config_json`,
|
||||
'$.emotion', 'neutral',
|
||||
'$.emotion_scale', 4
|
||||
)
|
||||
WHERE `id` = 'TTS_HuoshanDoubleStreamTTS'
|
||||
AND JSON_EXTRACT(`config_json`, '$.emotion') IS NULL
|
||||
AND JSON_EXTRACT(`config_json`, '$.emotion_scale') IS NULL;
|
||||
|
||||
-- 增加文档链接和备注
|
||||
UPDATE `ai_model_config` SET
|
||||
`doc_link` = 'https://console.volcengine.com/speech/service/10007',
|
||||
`remark` = '火山引擎语音合成服务配置说明:
|
||||
1. 访问 https://www.volcengine.com/ 注册并开通火山引擎账号
|
||||
2. 访问 https://console.volcengine.com/speech/service/10007 开通语音合成大模型,购买音色
|
||||
3. 在页面底部获取appid和access_token
|
||||
5. 资源ID固定为:volc.service_type.10029(大模型语音合成及混音)
|
||||
6. 链接复用:开启WebSocket连接复用,默认true减少链接损耗(注意:复用后设备处于聆听状态时空闲链接会占并发数)
|
||||
7. 语速:-50~100,可不填,正常默认值0,可填-50~100
|
||||
8. 音量:-50~100,可不填,正常默认值0,可填-50~100
|
||||
9. 音高:-12~12,可不填,正常默认值0,可填-12~12
|
||||
10. 多情感参数(当前仅部分音色支持设置情感):
|
||||
相关音色列表:https://www.volcengine.com/docs/6561/1257544
|
||||
- emotion_scale:情感强度,可选值为:1~5,默认值为4
|
||||
- emotion:情感类型,可选值为:neutral、happy、sad、angry、fearful、disgusted、surprised
|
||||
' WHERE `id` = 'TTS_HuoshanDoubleStreamTTS';
|
||||
@@ -424,6 +424,13 @@ databaseChangeLog:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202511131023.sql
|
||||
- changeSet:
|
||||
id: 202511221450
|
||||
author: RanChen
|
||||
changes:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202511221450.sql
|
||||
- changeSet:
|
||||
id: 202512031517
|
||||
author: rainv123
|
||||
changes:
|
||||
@@ -445,3 +452,31 @@ databaseChangeLog:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202512131453.sql
|
||||
- changeSet:
|
||||
id: 202512161529
|
||||
author: RanChen
|
||||
changes:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202512161529.sql
|
||||
- changeSet:
|
||||
id: 202512192245
|
||||
author: hrz
|
||||
changes:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202512192245.sql
|
||||
- changeSet:
|
||||
id: 202512221117
|
||||
author: RanChen
|
||||
changes:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202512221117.sql
|
||||
- changeSet:
|
||||
id: 202512301430
|
||||
author: RanChen
|
||||
changes:
|
||||
- sqlFile:
|
||||
encoding: utf8
|
||||
path: classpath:db/changelog/202512301430.sql
|
||||
|
||||
@@ -22,13 +22,18 @@
|
||||
created_at, updated_at
|
||||
</sql>
|
||||
|
||||
<delete id="deleteAudioByAgentId">
|
||||
DELETE FROM ai_agent_chat_audio
|
||||
WHERE id IN (
|
||||
SELECT audio_id
|
||||
FROM ai_agent_chat_history
|
||||
WHERE agent_id = #{agentId}
|
||||
)
|
||||
<select id="getAudioIdsByAgentId" resultType="java.lang.String">
|
||||
SELECT DISTINCT audio_id
|
||||
FROM ai_agent_chat_history
|
||||
WHERE agent_id = #{agentId} AND audio_id IS NOT NULL
|
||||
</select>
|
||||
|
||||
<delete id="deleteAudioByIds">
|
||||
DELETE FROM ai_agent_chat_audio
|
||||
WHERE id IN
|
||||
<foreach collection="audioIds" item="id" open="(" separator="," close=")">
|
||||
#{id}
|
||||
</foreach>
|
||||
</delete>
|
||||
|
||||
<update id="deleteAudioIdByAgentId">
|
||||
|
||||
@@ -235,7 +235,7 @@ function showAbout() {
|
||||
title: t('settings.aboutApp', { appName: import.meta.env.VITE_APP_TITLE }),
|
||||
content: t('settings.aboutContent', {
|
||||
appName: import.meta.env.VITE_APP_TITLE,
|
||||
version: '0.8.10'
|
||||
version: '0.8.11'
|
||||
}),
|
||||
showCancel: false,
|
||||
confirmText: t('common.confirm'),
|
||||
|
||||
|
Before Width: | Height: | Size: 1.9 KiB After Width: | Height: | Size: 1.9 KiB |
|
After Width: | Height: | Size: 6.8 KiB |
|
After Width: | Height: | Size: 7.2 KiB |
|
After Width: | Height: | Size: 7.0 KiB |
|
After Width: | Height: | Size: 1.9 KiB |
|
After Width: | Height: | Size: 1.8 KiB |
@@ -4,7 +4,7 @@
|
||||
<!-- 左侧元素 -->
|
||||
<div class="header-left" @click="goHome">
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-logo.png" class="logo-img" />
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-ai.png" class="brand-img" />
|
||||
<img loading="lazy" alt="" :src="xiaozhiAiIcon" class="brand-img" />
|
||||
</div>
|
||||
|
||||
<!-- 中间导航菜单 -->
|
||||
@@ -257,6 +257,24 @@ export default {
|
||||
return this.$t("language.zhCN");
|
||||
}
|
||||
},
|
||||
// 根据当前语言获取对应的xiaozhi-ai图标
|
||||
xiaozhiAiIcon() {
|
||||
const currentLang = this.currentLanguage;
|
||||
switch (currentLang) {
|
||||
case "zh_CN":
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
case "zh_TW":
|
||||
return require("@/assets/xiaozhi-ai_zh_TW.png");
|
||||
case "en":
|
||||
return require("@/assets/xiaozhi-ai_en.png");
|
||||
case "de":
|
||||
return require("@/assets/xiaozhi-ai_de.png");
|
||||
case "vi":
|
||||
return require("@/assets/xiaozhi-ai_vi.png");
|
||||
default:
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
}
|
||||
},
|
||||
// 用户菜单选项
|
||||
userMenuOptions() {
|
||||
return [
|
||||
|
||||
@@ -821,7 +821,7 @@ export default {
|
||||
'modelConfig.rag': 'RAG',
|
||||
'modelConfig.modelId': 'Modell-ID',
|
||||
'modelConfig.modelName': 'Modellname',
|
||||
'modelConfig.provider': 'Anbieter',
|
||||
'modelConfig.provider': 'Schnittstellentyp',
|
||||
'modelConfig.unknown': 'Unbekannt',
|
||||
'modelConfig.isEnabled': 'Aktiviert',
|
||||
'modelConfig.isDefault': 'Standard',
|
||||
|
||||
@@ -821,7 +821,7 @@ export default {
|
||||
'modelConfig.rag': 'RAG',
|
||||
'modelConfig.modelId': 'Model ID',
|
||||
'modelConfig.modelName': 'Model Name',
|
||||
'modelConfig.provider': 'Provider',
|
||||
'modelConfig.provider': 'Interface Type',
|
||||
'modelConfig.unknown': 'Unknown',
|
||||
'modelConfig.isEnabled': 'Enabled',
|
||||
'modelConfig.isDefault': 'Default',
|
||||
|
||||
@@ -821,7 +821,7 @@ export default {
|
||||
'modelConfig.rag': 'RAG',
|
||||
'modelConfig.modelId': 'ID mô hình',
|
||||
'modelConfig.modelName': 'Tên mô hình',
|
||||
'modelConfig.provider': 'Nhà cung cấp',
|
||||
'modelConfig.provider': 'Loại giao diện',
|
||||
'modelConfig.unknown': 'Không xác định',
|
||||
'modelConfig.isEnabled': 'Đã bật',
|
||||
'modelConfig.isDefault': 'Mặc định',
|
||||
|
||||
@@ -821,7 +821,7 @@ export default {
|
||||
'modelConfig.rag': '知识库',
|
||||
'modelConfig.modelId': '模型ID',
|
||||
'modelConfig.modelName': '模型名称',
|
||||
'modelConfig.provider': '提供商',
|
||||
'modelConfig.provider': '接口类型',
|
||||
'modelConfig.unknown': '未知',
|
||||
'modelConfig.isEnabled': '是否启用',
|
||||
'modelConfig.isDefault': '是否默认',
|
||||
|
||||
@@ -821,7 +821,7 @@ export default {
|
||||
'modelConfig.rag': '知識庫',
|
||||
'modelConfig.modelId': '模型ID',
|
||||
'modelConfig.modelName': '模型名稱',
|
||||
'modelConfig.provider': '提供商',
|
||||
'modelConfig.provider': '接口類型',
|
||||
'modelConfig.unknown': '未知',
|
||||
'modelConfig.isEnabled': '是否啟用',
|
||||
'modelConfig.isDefault': '是否默認',
|
||||
|
||||
@@ -35,7 +35,7 @@
|
||||
<el-table-column :label="$t('device.bindTime')" prop="bindTime" align="center"></el-table-column>
|
||||
<el-table-column :label="$t('device.lastConversation')" prop="lastConversation"
|
||||
align="center"></el-table-column>
|
||||
<el-table-column :label="$t('device.deviceStatus')" prop="deviceStatus" align="center">
|
||||
<el-table-column v-if="mqttServiceAvailable" :label="$t('device.deviceStatus')" prop="deviceStatus" align="center">
|
||||
<template slot-scope="scope">
|
||||
<el-tag v-if="scope.row.deviceStatus === 'online'" type="success">{{ $t('device.online') }}</el-tag>
|
||||
<el-tag v-else type="danger">{{ $t('device.offline') }}</el-tag>
|
||||
@@ -147,6 +147,7 @@ export default {
|
||||
loading: false,
|
||||
userApi: null,
|
||||
firmwareTypes: [],
|
||||
mqttServiceAvailable: false, // MQTT服务是否可用
|
||||
};
|
||||
},
|
||||
computed: {
|
||||
@@ -392,12 +393,21 @@ export default {
|
||||
|
||||
// 直接使用解析后的数据作为设备状态映射(不需要devices字段包装)
|
||||
if (statusData && typeof statusData === 'object') {
|
||||
// 成功获取到设备状态
|
||||
this.mqttServiceAvailable = true;
|
||||
// 更新设备状态
|
||||
this.updateDeviceStatusFromResponse(statusData);
|
||||
} else {
|
||||
// 数据格式不正确,MQTT服务不可用
|
||||
this.mqttServiceAvailable = false;
|
||||
}
|
||||
} catch (error) {
|
||||
// JSON解析失败,忽略状态更新
|
||||
// JSON解析失败,MQTT服务不可用
|
||||
this.mqttServiceAvailable = false;
|
||||
}
|
||||
} else {
|
||||
// 接口调用失败,MQTT服务不可用
|
||||
this.mqttServiceAvailable = false;
|
||||
}
|
||||
});
|
||||
},
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
gap: 10px;
|
||||
">
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-logo.png" style="width: 45px; height: 45px" />
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-ai.png" style="height: 18px" />
|
||||
<img loading="lazy" alt="" :src="xiaozhiAiIcon" style="height: 18px" />
|
||||
</div>
|
||||
</el-header>
|
||||
<div class="login-person">
|
||||
@@ -192,6 +192,24 @@ export default {
|
||||
return this.$t("language.zhCN");
|
||||
}
|
||||
},
|
||||
// 根据当前语言获取对应的xiaozhi-ai图标
|
||||
xiaozhiAiIcon() {
|
||||
const currentLang = this.currentLanguage;
|
||||
switch (currentLang) {
|
||||
case "zh_CN":
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
case "zh_TW":
|
||||
return require("@/assets/xiaozhi-ai_zh_TW.png");
|
||||
case "en":
|
||||
return require("@/assets/xiaozhi-ai_en.png");
|
||||
case "de":
|
||||
return require("@/assets/xiaozhi-ai_de.png");
|
||||
case "vi":
|
||||
return require("@/assets/xiaozhi-ai_vi.png");
|
||||
default:
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
}
|
||||
},
|
||||
},
|
||||
data() {
|
||||
return {
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
<el-header>
|
||||
<div style="display: flex;align-items: center;margin-top: 15px;margin-left: 10px;gap: 10px;">
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-logo.png" style="width: 45px;height: 45px;" />
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-ai.png" style="height: 18px;" />
|
||||
<img loading="lazy" alt="" :src="xiaozhiAiIcon" style="height: 18px;" />
|
||||
</div>
|
||||
</el-header>
|
||||
<div class="login-person">
|
||||
@@ -108,7 +108,7 @@
|
||||
<div style="font-size: 14px;color: #979db1;">
|
||||
{{ $t('register.agreeTo') }}
|
||||
<div style="display: inline-block;color: #5778FF;cursor: pointer;">{{ $t('register.userAgreement') }}</div>
|
||||
{{ $t('register.and') }}
|
||||
{{ $t('login.and') }}
|
||||
<div style="display: inline-block;color: #5778FF;cursor: pointer;">{{ $t('register.privacyPolicy') }}</div>
|
||||
</div>
|
||||
</div>
|
||||
@@ -127,6 +127,7 @@ import Api from '@/apis/api';
|
||||
import VersionFooter from '@/components/VersionFooter.vue';
|
||||
import { getUUID, goToPage, showDanger, showSuccess, sm2Encrypt, validateMobile } from '@/utils';
|
||||
import { mapState } from 'vuex';
|
||||
import i18n from '@/i18n';
|
||||
|
||||
// 导入语言切换功能
|
||||
|
||||
@@ -142,6 +143,28 @@ export default {
|
||||
mobileAreaList: state => state.pubConfig.mobileAreaList,
|
||||
sm2PublicKey: state => state.pubConfig.sm2PublicKey,
|
||||
}),
|
||||
// 获取当前语言
|
||||
currentLanguage() {
|
||||
return i18n.locale || "zh_CN";
|
||||
},
|
||||
// 根据当前语言获取对应的xiaozhi-ai图标
|
||||
xiaozhiAiIcon() {
|
||||
const currentLang = this.currentLanguage;
|
||||
switch (currentLang) {
|
||||
case "zh_CN":
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
case "zh_TW":
|
||||
return require("@/assets/xiaozhi-ai_zh_TW.png");
|
||||
case "en":
|
||||
return require("@/assets/xiaozhi-ai_en.png");
|
||||
case "de":
|
||||
return require("@/assets/xiaozhi-ai_de.png");
|
||||
case "vi":
|
||||
return require("@/assets/xiaozhi-ai_vi.png");
|
||||
default:
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
}
|
||||
},
|
||||
canSendMobileCaptcha() {
|
||||
return this.countdown === 0 && validateMobile(this.form.mobile, this.form.areaCode);
|
||||
}
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
<el-header>
|
||||
<div style="display: flex;align-items: center;margin-top: 15px;margin-left: 10px;gap: 10px;">
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-logo.png" style="width: 45px;height: 45px;" />
|
||||
<img loading="lazy" alt="" src="@/assets/xiaozhi-ai.png" style="height: 18px;" />
|
||||
<img loading="lazy" alt="" :src="xiaozhiAiIcon" style="height: 18px;" />
|
||||
</div>
|
||||
</el-header>
|
||||
<div class="login-person">
|
||||
@@ -83,7 +83,7 @@
|
||||
<div style="font-size: 14px;color: #979db1;">
|
||||
{{ $t('retrievePassword.agreeTo') }}
|
||||
<div style="display: inline-block;color: #5778FF;cursor: pointer;">{{ $t('register.userAgreement') }}</div>
|
||||
{{ $t('register.and') }}
|
||||
{{ $t('login.and') }}
|
||||
<div style="display: inline-block;color: #5778FF;cursor: pointer;">{{ $t('register.privacyPolicy') }}</div>
|
||||
</div>
|
||||
</div>
|
||||
@@ -103,6 +103,7 @@ import Api from '@/apis/api';
|
||||
import VersionFooter from '@/components/VersionFooter.vue';
|
||||
import { getUUID, goToPage, showDanger, showSuccess, validateMobile, sm2Encrypt } from '@/utils';
|
||||
import { mapState } from 'vuex';
|
||||
import i18n from '@/i18n';
|
||||
|
||||
// 导入语言切换功能
|
||||
import { changeLanguage } from '@/i18n';
|
||||
@@ -118,6 +119,28 @@ export default {
|
||||
mobileAreaList: state => state.pubConfig.mobileAreaList,
|
||||
sm2PublicKey: state => state.pubConfig.sm2PublicKey
|
||||
}),
|
||||
// 获取当前语言
|
||||
currentLanguage() {
|
||||
return i18n.locale || "zh_CN";
|
||||
},
|
||||
// 根据当前语言获取对应的xiaozhi-ai图标
|
||||
xiaozhiAiIcon() {
|
||||
const currentLang = this.currentLanguage;
|
||||
switch (currentLang) {
|
||||
case "zh_CN":
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
case "zh_TW":
|
||||
return require("@/assets/xiaozhi-ai_zh_TW.png");
|
||||
case "en":
|
||||
return require("@/assets/xiaozhi-ai_en.png");
|
||||
case "de":
|
||||
return require("@/assets/xiaozhi-ai_de.png");
|
||||
case "vi":
|
||||
return require("@/assets/xiaozhi-ai_vi.png");
|
||||
default:
|
||||
return require("@/assets/xiaozhi-ai.png");
|
||||
}
|
||||
},
|
||||
canSendMobileCaptcha() {
|
||||
return this.countdown === 0 && validateMobile(this.form.mobile, this.form.areaCode);
|
||||
}
|
||||
|
||||
@@ -69,6 +69,9 @@ enable_greeting: true
|
||||
enable_stop_tts_notify: false
|
||||
# 说完话是否开启提示音,音效地址
|
||||
stop_tts_notify_voice: "config/assets/tts_notify.mp3"
|
||||
# 是否启用WebSocket心跳保活机制
|
||||
enable_websocket_ping: false
|
||||
|
||||
|
||||
# TTS音频发送延迟配置
|
||||
# tts_audio_send_delay: 控制音频包发送间隔
|
||||
@@ -343,6 +346,13 @@ ASR:
|
||||
# 热词、替换词使用流程:https://www.volcengine.com/docs/6561/155738
|
||||
boosting_table_name: (选填)你的热词文件名称
|
||||
correct_table_name: (选填)你的替换词文件名称
|
||||
# 是否开启多语种识别模式
|
||||
enable_multilingual: False
|
||||
# 多语种识别当该键为空时,该模型支持中英文、上海话、闽南语,四川、陕西、粤语识别。当将其设置为特定键时,它可以识别指定语言。
|
||||
# 详细语言列表参考 https://www.volcengine.com/docs/6561/1354869
|
||||
# language: zh-cn
|
||||
# 静音判定时长(ms),默认200ms
|
||||
end_window_size: 200
|
||||
output_dir: tmp/
|
||||
TencentASR:
|
||||
# token申请地址:https://console.cloud.tencent.com/cam/capi
|
||||
@@ -473,7 +483,32 @@ ASR:
|
||||
accent: mandarin # 方言,mandarin:普通话
|
||||
# 调整音频处理参数以提高长语音识别质量
|
||||
output_dir: tmp/
|
||||
|
||||
AliyunBLStreamASR:
|
||||
# 阿里百炼Paraformer实时语音识别服务
|
||||
# WebSocket实时流式语音识别,支持多语言、热词定制、语义断句等高级功能
|
||||
# 平台地址:https://bailian.console.aliyun.com/
|
||||
# API Key地址:https://bailian.console.aliyun.com/#/api-key
|
||||
# 文档地址:https://help.aliyun.com/zh/model-studio/websocket-for-paraformer-real-time-service
|
||||
# 支持模型:paraformer-realtime-v2(推荐), paraformer-realtime-8k-v2, paraformer-realtime-v1, paraformer-realtime-8k-v1
|
||||
type: aliyunbl_stream
|
||||
# 必填参数
|
||||
api_key: 你的阿里云百炼API密钥
|
||||
# 模型选择,推荐使用v2版本
|
||||
model: paraformer-realtime-v2
|
||||
# 音频格式和采样率
|
||||
format: pcm
|
||||
sample_rate: 16000 # v2支持任意采样率,v1仅支持16000,8k版本仅支持8000
|
||||
# 可选参数
|
||||
disfluency_removal_enabled: false # 是否过滤语气词(如"嗯"、"啊"等)
|
||||
semantic_punctuation_enabled: false # 语义断句(true:会议场景,准确;false:VAD断句,交互场景,低延迟)
|
||||
max_sentence_silence: 200 # VAD断句静音时长阈值(ms),范围200-6000,仅VAD断句时生效
|
||||
multi_threshold_mode_enabled: false # 防止VAD断句切割过长,仅VAD断句时生效
|
||||
punctuation_prediction_enabled: true # 是否自动添加标点符号
|
||||
inverse_text_normalization_enabled: true # 是否开启ITN(中文数字转阿拉伯数字)
|
||||
# 热词定制文档地址:https://help.aliyun.com/zh/model-studio/custom-hot-words?
|
||||
# vocabulary_id: vocab-xxx-24ee19fa8cfb4d52902170a0xxxxxxxx # 热词ID(可选)
|
||||
# language_hints: ["zh", "en"] # 指定语言(可选),支持zh、en、ja、yue、ko、de、fr、ru
|
||||
output_dir: tmp/
|
||||
VAD:
|
||||
SileroVAD:
|
||||
type: silero
|
||||
@@ -677,9 +712,15 @@ TTS:
|
||||
access_token: 你的火山引擎语音合成服务access_token
|
||||
resource_id: volc.service_type.10029
|
||||
speaker: zh_female_wanwanxiaohe_moon_bigtts
|
||||
# 开启WebSocket连接复用,默认复用(注意:复用后设备处于聆听状态时空闲链接会占并发数)
|
||||
enable_ws_reuse: True
|
||||
speech_rate: 0
|
||||
loudness_rate: 0
|
||||
pitch: 0
|
||||
# 多情感音色参数,注意:当前仅部分音色支持设置情感。
|
||||
# 相关音色列表:https://www.volcengine.com/docs/6561/1257544
|
||||
emotion: "neutral" # 情感类型,可选值为:neutral、happy、sad、angry、fearful、disgusted、surprised
|
||||
emotion_scale: 4 # 情感强度,可选值为:1~5,默认值为4
|
||||
CosyVoiceSiliconflow:
|
||||
type: siliconflow
|
||||
# 硅基流动TTS
|
||||
|
||||
@@ -5,7 +5,7 @@ from config.config_loader import load_config
|
||||
from config.settings import check_config_file
|
||||
from datetime import datetime
|
||||
|
||||
SERVER_VERSION = "0.8.10"
|
||||
SERVER_VERSION = "0.8.11"
|
||||
_logger_initialized = False
|
||||
|
||||
|
||||
|
||||
@@ -53,6 +53,7 @@ class ManageApiClient:
|
||||
async def _ensure_async_client(cls):
|
||||
"""确保异步客户端已创建(为每个事件循环创建独立的客户端)"""
|
||||
import asyncio
|
||||
|
||||
try:
|
||||
loop = asyncio.get_running_loop()
|
||||
loop_id = id(loop)
|
||||
@@ -115,6 +116,7 @@ class ManageApiClient:
|
||||
async def _execute_async_request(cls, method: str, endpoint: str, **kwargs) -> Dict:
|
||||
"""带重试机制的异步请求执行器"""
|
||||
import asyncio
|
||||
|
||||
retry_count = 0
|
||||
|
||||
while retry_count <= cls.max_retries:
|
||||
@@ -138,6 +140,7 @@ class ManageApiClient:
|
||||
def safe_close(cls):
|
||||
"""安全关闭所有异步连接池"""
|
||||
import asyncio
|
||||
|
||||
for client in list(cls._async_clients.values()):
|
||||
try:
|
||||
asyncio.run(client.aclose())
|
||||
@@ -149,7 +152,9 @@ class ManageApiClient:
|
||||
|
||||
async def get_server_config() -> Optional[Dict]:
|
||||
"""获取服务器基础配置"""
|
||||
return await ManageApiClient._instance._execute_async_request("POST", "/config/server-base")
|
||||
return await ManageApiClient._instance._execute_async_request(
|
||||
"POST", "/config/server-base"
|
||||
)
|
||||
|
||||
|
||||
async def get_agent_models(
|
||||
@@ -167,17 +172,15 @@ async def get_agent_models(
|
||||
)
|
||||
|
||||
|
||||
async def save_mem_local_short(mac_address: str, short_momery: str) -> Optional[Dict]:
|
||||
async def generate_and_save_chat_summary(session_id: str) -> Optional[Dict]:
|
||||
"""生成并保存聊天记录总结"""
|
||||
try:
|
||||
return await ManageApiClient._instance._execute_async_request(
|
||||
"PUT",
|
||||
f"/agent/saveMemory/" + mac_address,
|
||||
json={
|
||||
"summaryMemory": short_momery,
|
||||
},
|
||||
"POST",
|
||||
f"/agent/chat-summary/{session_id}/save",
|
||||
)
|
||||
except Exception as e:
|
||||
print(f"存储短期记忆到服务器失败: {e}")
|
||||
print(f"生成并保存聊天记录总结失败: {e}")
|
||||
return None
|
||||
|
||||
|
||||
|
||||
@@ -10,7 +10,15 @@ class BaseHandler:
|
||||
def _add_cors_headers(self, response):
|
||||
"""添加CORS头信息"""
|
||||
response.headers["Access-Control-Allow-Headers"] = (
|
||||
"client-id, content-type, device-id"
|
||||
"client-id, content-type, device-id, authorization"
|
||||
)
|
||||
response.headers["Access-Control-Allow-Credentials"] = "true"
|
||||
response.headers["Access-Control-Allow-Origin"] = "*"
|
||||
|
||||
async def handle_options(self, request):
|
||||
"""处理OPTIONS请求,添加CORS头信息"""
|
||||
response = web.Response(body=b"", content_type="text/plain")
|
||||
self._add_cors_headers(response)
|
||||
# 添加允许的方法
|
||||
response.headers["Access-Control-Allow-Methods"] = "GET, POST, OPTIONS"
|
||||
return response
|
||||
|
||||
@@ -3,15 +3,46 @@ import time
|
||||
import base64
|
||||
import hashlib
|
||||
import hmac
|
||||
import os
|
||||
import re
|
||||
import glob
|
||||
from typing import Dict, List, Tuple
|
||||
from aiohttp import web
|
||||
|
||||
from core.auth import AuthManager
|
||||
from core.utils.util import get_local_ip
|
||||
from core.utils.util import get_local_ip, get_vision_url
|
||||
from core.api.base_handler import BaseHandler
|
||||
|
||||
TAG = __name__
|
||||
|
||||
|
||||
def _safe_basename(filename: str) -> str:
|
||||
# Prevent directory traversal
|
||||
return os.path.basename(filename)
|
||||
|
||||
|
||||
def _parse_version(ver: str) -> Tuple[int, ...]:
|
||||
# conservative parser: split by non-digit, keep numeric parts
|
||||
parts = re.findall(r"\d+", ver)
|
||||
return tuple(int(p) for p in parts) if parts else (0,)
|
||||
|
||||
|
||||
def _is_higher_version(a: str, b: str) -> bool:
|
||||
"""Return True if version string a > b (semver-like numeric compare)."""
|
||||
ta = _parse_version(a)
|
||||
tb = _parse_version(b)
|
||||
# compare tuple lexicographically, but allow different lengths
|
||||
maxlen = max(len(ta), len(tb))
|
||||
for i in range(maxlen):
|
||||
ai = ta[i] if i < len(ta) else 0
|
||||
bi = tb[i] if i < len(tb) else 0
|
||||
if ai > bi:
|
||||
return True
|
||||
if ai < bi:
|
||||
return False
|
||||
return False
|
||||
|
||||
|
||||
class OTAHandler(BaseHandler):
|
||||
def __init__(self, config: dict):
|
||||
super().__init__(config)
|
||||
@@ -23,6 +54,54 @@ class OTAHandler(BaseHandler):
|
||||
expire_seconds = auth_config.get("expire_seconds")
|
||||
self.auth = AuthManager(secret_key=secret_key, expire_seconds=expire_seconds)
|
||||
|
||||
# firmware storage
|
||||
self.bin_dir = os.path.join(os.getcwd(), "data", "bin")
|
||||
# cache structure: { 'updated_at': timestamp, 'ttl': seconds, 'files_by_model': { model: [(version, filename), ...] } }
|
||||
self._bin_cache: Dict = {
|
||||
"updated_at": 0,
|
||||
"ttl": config.get("firmware_cache_ttl", 30),
|
||||
"files_by_model": {},
|
||||
}
|
||||
|
||||
def _refresh_bin_cache_if_needed(self):
|
||||
now = int(time.time())
|
||||
ttl = int(self._bin_cache.get("ttl", 30))
|
||||
if now - int(
|
||||
self._bin_cache.get("updated_at", 0)
|
||||
) < ttl and self._bin_cache.get("files_by_model"):
|
||||
return
|
||||
|
||||
files_by_model: Dict[str, List[Tuple[str, str]]] = {}
|
||||
try:
|
||||
if not os.path.isdir(self.bin_dir):
|
||||
os.makedirs(self.bin_dir, exist_ok=True)
|
||||
|
||||
# match files like model_1.2.3.bin (allow dots, dashes, underscores in model and version)
|
||||
pattern = os.path.join(self.bin_dir, "*.bin")
|
||||
for path in glob.glob(pattern):
|
||||
fname = os.path.basename(path)
|
||||
# filename format: {model}_{version}.bin
|
||||
m = re.match(r"^(.+?)_([0-9][A-Za-z0-9\.\-_]*)\.bin$", fname)
|
||||
if not m:
|
||||
# skip files not conforming to naming rule
|
||||
continue
|
||||
model = m.group(1)
|
||||
version = m.group(2)
|
||||
files_by_model.setdefault(model, []).append((version, fname))
|
||||
|
||||
# sort versions for each model descending
|
||||
for model, items in files_by_model.items():
|
||||
items.sort(key=lambda it: _parse_version(it[0]), reverse=True)
|
||||
|
||||
self._bin_cache["files_by_model"] = files_by_model
|
||||
self._bin_cache["updated_at"] = now
|
||||
self.logger.bind(tag=TAG).info(
|
||||
f"Firmware cache refreshed: {len(files_by_model)} models"
|
||||
)
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"刷新固件缓存失败: {e}")
|
||||
# keep previous cache if any
|
||||
|
||||
def generate_password_signature(self, content: str, secret_key: str) -> str:
|
||||
"""生成MQTT密码签名
|
||||
|
||||
@@ -62,7 +141,14 @@ class OTAHandler(BaseHandler):
|
||||
return f"ws://{local_ip}:{port}/xiaozhi/v1/"
|
||||
|
||||
async def handle_post(self, request):
|
||||
"""处理 OTA POST 请求"""
|
||||
"""处理 OTA POST 请求
|
||||
|
||||
This handler will:
|
||||
- read device id/client id (as before)
|
||||
- attempt to determine device model and current firmware version (prefer headers, fallback to body)
|
||||
- check data/bin for newer firmware for that model
|
||||
- if found a newer firmware, set firmware.url to the download endpoint
|
||||
"""
|
||||
try:
|
||||
data = await request.text()
|
||||
self.logger.bind(tag=TAG).debug(f"OTA请求方法: {request.method}")
|
||||
@@ -81,33 +167,76 @@ class OTAHandler(BaseHandler):
|
||||
else:
|
||||
raise Exception("OTA请求ClientID为空")
|
||||
|
||||
data_json = json.loads(data)
|
||||
data_json = {}
|
||||
try:
|
||||
data_json = json.loads(data) if data else {}
|
||||
except Exception:
|
||||
data_json = {}
|
||||
|
||||
server_config = self.config["server"]
|
||||
port = int(server_config.get("port", 8000))
|
||||
# Distinguish ports:
|
||||
# - websocket_port is used to construct websocket URL (server["port"])
|
||||
# - http_port is used to construct OTA download URLs (server["http_port"])
|
||||
websocket_port = int(server_config.get("port", 8000))
|
||||
http_port = int(server_config.get("http_port", 8003))
|
||||
local_ip = get_local_ip()
|
||||
|
||||
# Determine device model (prefer headers)
|
||||
device_model = ""
|
||||
# header candidates
|
||||
for h in ("device-model", "device_model", "model"):
|
||||
if h in request.headers:
|
||||
device_model = request.headers.get(h, "").strip()
|
||||
break
|
||||
# body fallback
|
||||
if not device_model:
|
||||
try:
|
||||
if "board" in data_json and isinstance(data_json["board"], dict):
|
||||
device_model = data_json["board"].get("type", "")
|
||||
elif "model" in data_json:
|
||||
device_model = data_json.get("model", "")
|
||||
except Exception:
|
||||
device_model = ""
|
||||
if not device_model:
|
||||
device_model = "default"
|
||||
|
||||
# Determine device current version (prefer headers)
|
||||
device_version = ""
|
||||
for h in (
|
||||
"device-version",
|
||||
"device_version",
|
||||
"firmware-version",
|
||||
"app-version",
|
||||
"application-version",
|
||||
):
|
||||
if h in request.headers:
|
||||
device_version = request.headers.get(h, "").strip()
|
||||
break
|
||||
if not device_version:
|
||||
try:
|
||||
device_version = data_json.get("application", {}).get("version", "")
|
||||
except Exception:
|
||||
device_version = ""
|
||||
if not device_version:
|
||||
device_version = "0.0.0"
|
||||
|
||||
return_json = {
|
||||
"server_time": {
|
||||
"timestamp": int(round(time.time() * 1000)),
|
||||
"timezone_offset": server_config.get("timezone_offset", 8) * 60,
|
||||
},
|
||||
"firmware": {
|
||||
"version": data_json["application"].get("version", "1.0.0"),
|
||||
"version": device_version,
|
||||
"url": "",
|
||||
},
|
||||
}
|
||||
|
||||
# existing mqtt/websocket logic (unchanged)
|
||||
mqtt_gateway_endpoint = server_config.get("mqtt_gateway")
|
||||
|
||||
if mqtt_gateway_endpoint: # 如果配置了非空字符串
|
||||
# 尝试从请求数据中获取设备型号
|
||||
device_model = "default"
|
||||
# 尝试从请求数据中获取设备型号(已解析 above)
|
||||
try:
|
||||
if "device" in data_json and isinstance(data_json["device"], dict):
|
||||
device_model = data_json["device"].get("model", "default")
|
||||
elif "model" in data_json:
|
||||
device_model = data_json["model"]
|
||||
group_id = f"GID_{device_model}".replace(":", "_").replace(" ", "_")
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"获取设备型号失败: {e}")
|
||||
@@ -159,20 +288,61 @@ class OTAHandler(BaseHandler):
|
||||
token = self.auth.generate_token(client_id, device_id)
|
||||
else:
|
||||
token = self.auth.generate_token(client_id, device_id)
|
||||
# NOTE: use websocket_port here
|
||||
return_json["websocket"] = {
|
||||
"url": self._get_websocket_url(local_ip, port),
|
||||
"url": self._get_websocket_url(local_ip, websocket_port),
|
||||
"token": token,
|
||||
}
|
||||
self.logger.bind(tag=TAG).info(
|
||||
f"未配置MQTT网关,为设备 {device_id} 下发WebSocket配置"
|
||||
)
|
||||
self.logger.bind(tag=TAG).info(f"{return_json}")
|
||||
|
||||
# Now check firmware files for updates
|
||||
try:
|
||||
self._refresh_bin_cache_if_needed()
|
||||
files_by_model = self._bin_cache.get("files_by_model", {})
|
||||
candidates = files_by_model.get(device_model, [])
|
||||
|
||||
self.logger.bind(tag=TAG).info(
|
||||
f"查找型号 {device_model} 的固件,找到 {len(candidates)} 个候选"
|
||||
)
|
||||
|
||||
chosen_url = ""
|
||||
chosen_version = device_version
|
||||
|
||||
# candidates are sorted descending by version
|
||||
for ver, fname in candidates:
|
||||
if _is_higher_version(ver, device_version):
|
||||
# build download url (only allow download via our download endpoint)
|
||||
chosen_version = ver
|
||||
# Use get_vision_url to get the base URL and replace the path
|
||||
vision_url = get_vision_url(self.config)
|
||||
# Replace the path from "/mcp/vision/explain" to "/xiaozhi/ota/download/{fname}"
|
||||
chosen_url = vision_url.replace(
|
||||
"/mcp/vision/explain", f"/xiaozhi/ota/download/{fname}"
|
||||
)
|
||||
break
|
||||
|
||||
if chosen_url:
|
||||
return_json["firmware"]["version"] = chosen_version
|
||||
return_json["firmware"]["url"] = chosen_url
|
||||
self.logger.bind(tag=TAG).info(
|
||||
f"为设备 {device_id} 下发固件 {chosen_version} [如果地址前缀有误,请检查配置文件中的server.vision_explain]-> {chosen_url} "
|
||||
)
|
||||
else:
|
||||
self.logger.bind(tag=TAG).info(
|
||||
f"设备 {device_id} 固件已是最新: {device_version}"
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"检查固件版本时出错: {e}")
|
||||
|
||||
response = web.Response(
|
||||
text=json.dumps(return_json, separators=(",", ":")),
|
||||
content_type="application/json",
|
||||
)
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"OTA POST处理异常: {e}")
|
||||
return_json = {"success": False, "message": "request error."}
|
||||
response = web.Response(
|
||||
text=json.dumps(return_json, separators=(",", ":")),
|
||||
@@ -187,8 +357,9 @@ class OTAHandler(BaseHandler):
|
||||
try:
|
||||
server_config = self.config["server"]
|
||||
local_ip = get_local_ip()
|
||||
port = int(server_config.get("port", 8000))
|
||||
websocket_url = self._get_websocket_url(local_ip, port)
|
||||
# use websocket port for websocket URL
|
||||
websocket_port = int(server_config.get("port", 8000))
|
||||
websocket_url = self._get_websocket_url(local_ip, websocket_port)
|
||||
message = f"OTA接口运行正常,向设备发送的websocket地址是:{websocket_url}"
|
||||
response = web.Response(text=message, content_type="text/plain")
|
||||
except Exception as e:
|
||||
@@ -197,3 +368,48 @@ class OTAHandler(BaseHandler):
|
||||
finally:
|
||||
self._add_cors_headers(response)
|
||||
return response
|
||||
|
||||
async def handle_download(self, request):
|
||||
"""
|
||||
下载固件接口
|
||||
URL: /xiaozhi/ota/download/{filename}
|
||||
- 只允许下载 data/bin 目录下的 .bin 文件
|
||||
- filename 必须是 basename 且匹配安全的模式
|
||||
"""
|
||||
try:
|
||||
fname = request.match_info.get("filename", "")
|
||||
if not fname:
|
||||
raise web.HTTPBadRequest(text="filename required")
|
||||
|
||||
# sanitize
|
||||
fname = _safe_basename(fname)
|
||||
# pattern: allow letters, numbers, dot, underscore, dash
|
||||
if not re.match(r"^[A-Za-z0-9\.\-_]+\.bin$", fname):
|
||||
raise web.HTTPBadRequest(text="invalid filename")
|
||||
|
||||
file_path = os.path.join(self.bin_dir, fname)
|
||||
# ensure realpath is under bin_dir
|
||||
file_real = os.path.realpath(file_path)
|
||||
bin_dir_real = os.path.realpath(self.bin_dir)
|
||||
if (
|
||||
not file_real.startswith(bin_dir_real + os.sep)
|
||||
and file_real != bin_dir_real
|
||||
):
|
||||
raise web.HTTPForbidden(text="forbidden")
|
||||
|
||||
if not os.path.isfile(file_real):
|
||||
raise web.HTTPNotFound(text="file not found")
|
||||
|
||||
# use FileResponse to stream file
|
||||
resp = web.FileResponse(path=file_real)
|
||||
except web.HTTPError as e:
|
||||
resp = e
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"固件下载异常: {e}")
|
||||
resp = web.Response(text="download error", status=500)
|
||||
finally:
|
||||
try:
|
||||
self._add_cors_headers(resp)
|
||||
except Exception:
|
||||
pass
|
||||
return resp
|
||||
|
||||
@@ -2,6 +2,7 @@ import json
|
||||
import copy
|
||||
from aiohttp import web
|
||||
from config.logger import setup_logging
|
||||
from core.api.base_handler import BaseHandler
|
||||
from core.utils.util import get_vision_url, is_valid_image_file
|
||||
from core.utils.vllm import create_instance
|
||||
from config.config_loader import get_private_config_from_api
|
||||
@@ -16,10 +17,9 @@ TAG = __name__
|
||||
MAX_FILE_SIZE = 5 * 1024 * 1024
|
||||
|
||||
|
||||
class VisionHandler:
|
||||
class VisionHandler(BaseHandler):
|
||||
def __init__(self, config: dict):
|
||||
self.config = config
|
||||
self.logger = setup_logging()
|
||||
super().__init__(config)
|
||||
# 初始化认证工具
|
||||
self.auth = AuthToken(config["server"]["auth_key"])
|
||||
|
||||
@@ -172,11 +172,3 @@ class VisionHandler:
|
||||
finally:
|
||||
self._add_cors_headers(response)
|
||||
return response
|
||||
|
||||
def _add_cors_headers(self, response):
|
||||
"""添加CORS头信息"""
|
||||
response.headers["Access-Control-Allow-Headers"] = (
|
||||
"client-id, content-type, device-id"
|
||||
)
|
||||
response.headers["Access-Control-Allow-Credentials"] = "true"
|
||||
response.headers["Access-Control-Allow-Origin"] = "*"
|
||||
|
||||
@@ -244,7 +244,9 @@ class ConnectionHandler:
|
||||
loop = asyncio.new_event_loop()
|
||||
asyncio.set_event_loop(loop)
|
||||
loop.run_until_complete(
|
||||
self.memory.save_memory(self.dialogue.dialogue)
|
||||
self.memory.save_memory(
|
||||
self.dialogue.dialogue, self.session_id
|
||||
)
|
||||
)
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"保存记忆失败: {e}")
|
||||
@@ -434,6 +436,7 @@ class ConnectionHandler:
|
||||
self.tts.open_audio_channels(self), self.loop
|
||||
)
|
||||
if self.need_bind:
|
||||
self.bind_completed_event.set()
|
||||
return
|
||||
self.selected_module_str = build_module_string(
|
||||
self.config.get("selected_module", {})
|
||||
@@ -574,16 +577,13 @@ class ConnectionHandler:
|
||||
self.bind_completed_event.set()
|
||||
except DeviceNotFoundException as e:
|
||||
self.need_bind = True
|
||||
self.bind_completed_event.set() # 状态已确定,设置事件
|
||||
private_config = {}
|
||||
except DeviceBindException as e:
|
||||
self.need_bind = True
|
||||
self.bind_code = e.bind_code
|
||||
self.bind_completed_event.set() # 状态已确定,设置事件
|
||||
private_config = {}
|
||||
except Exception as e:
|
||||
self.need_bind = True
|
||||
self.bind_completed_event.set() # 状态已确定,设置事件
|
||||
self.logger.bind(tag=TAG).error(f"异步获取差异化配置失败: {e}")
|
||||
private_config = {}
|
||||
|
||||
|
||||
@@ -7,6 +7,10 @@ from core.providers.tts.dto.dto import SentenceType
|
||||
from core.utils.audioRateController import AudioRateController
|
||||
|
||||
TAG = __name__
|
||||
# 音频帧时长(毫秒)
|
||||
AUDIO_FRAME_DURATION = 60
|
||||
# 预缓冲包数量,直接发送以减少延迟
|
||||
PRE_BUFFER_COUNT = 5
|
||||
|
||||
|
||||
async def sendAudioMessage(conn, sentenceType, audios, text):
|
||||
@@ -45,7 +49,7 @@ async def sendAudioMessage(conn, sentenceType, audios, text):
|
||||
|
||||
async def _wait_for_audio_completion(conn):
|
||||
"""
|
||||
等待音频队列清空
|
||||
等待音频队列清空并等待预缓冲包播放完成
|
||||
|
||||
Args:
|
||||
conn: 连接对象
|
||||
@@ -56,6 +60,13 @@ async def _wait_for_audio_completion(conn):
|
||||
f"等待音频发送完成,队列中还有 {len(rate_controller.queue)} 个包"
|
||||
)
|
||||
await rate_controller.queue_empty_event.wait()
|
||||
|
||||
# 等待预缓冲包播放完成
|
||||
# 前N个包直接发送,增加2个网络抖动包,需要额外等待它们在客户端播放完成
|
||||
frame_duration_ms = rate_controller.frame_duration
|
||||
pre_buffer_playback_time = (PRE_BUFFER_COUNT + 2) * frame_duration_ms / 1000.0
|
||||
await asyncio.sleep(pre_buffer_playback_time)
|
||||
|
||||
conn.logger.bind(tag=TAG).debug("音频发送完成")
|
||||
|
||||
|
||||
@@ -81,14 +92,14 @@ async def _send_to_mqtt_gateway(conn, opus_packet, timestamp, sequence):
|
||||
await conn.websocket.send(complete_packet)
|
||||
|
||||
|
||||
async def sendAudio(conn, audios, frame_duration=60):
|
||||
async def sendAudio(conn, audios, frame_duration=AUDIO_FRAME_DURATION):
|
||||
"""
|
||||
发送音频包,使用 AudioRateController 进行精确的流量控制
|
||||
|
||||
Args:
|
||||
conn: 连接对象
|
||||
audios: 单个opus包(bytes) 或 opus包列表
|
||||
frame_duration: 帧时长(毫秒),默认60ms
|
||||
frame_duration: 帧时长(毫秒),默认使用全局常量AUDIO_FRAME_DURATION
|
||||
"""
|
||||
if audios is None or len(audios) == 0:
|
||||
return
|
||||
@@ -187,16 +198,14 @@ async def _send_audio_with_rate_control(
|
||||
flow_control: 流控状态
|
||||
send_delay: 固定延迟(秒),-1表示使用动态流控
|
||||
"""
|
||||
pre_buffer_count = 5
|
||||
|
||||
for packet in audio_list:
|
||||
if conn.client_abort:
|
||||
return
|
||||
|
||||
conn.last_activity_time = time.time() * 1000
|
||||
|
||||
# 预缓冲:前5个包直接发送
|
||||
if flow_control["packet_count"] < pre_buffer_count:
|
||||
# 预缓冲:前N个包直接发送
|
||||
if flow_control["packet_count"] < PRE_BUFFER_COUNT:
|
||||
await _do_send_audio(conn, packet, flow_control)
|
||||
conn.client_is_speaking = True
|
||||
elif send_delay > 0:
|
||||
|
||||
@@ -69,6 +69,7 @@ class ListenTextMessageHandler(TextMessageHandler):
|
||||
enqueue_asr_report(conn, "嘿,你好呀", [])
|
||||
await startToChat(conn, "嘿,你好呀")
|
||||
else:
|
||||
conn.just_woken_up = True
|
||||
# 上报纯文字数据(复用ASR上报功能,但不提供音频数据)
|
||||
enqueue_asr_report(conn, original_text, [])
|
||||
# 否则需要LLM对文字内容进行答复
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
import json
|
||||
import time
|
||||
from typing import Dict, Any
|
||||
|
||||
from core.handle.textMessageHandler import TextMessageHandler
|
||||
from core.handle.textMessageType import TextMessageType
|
||||
|
||||
TAG = __name__
|
||||
|
||||
|
||||
class PingMessageHandler(TextMessageHandler):
|
||||
"""Ping消息处理器,用于保持WebSocket连接"""
|
||||
|
||||
@property
|
||||
def message_type(self) -> TextMessageType:
|
||||
return TextMessageType.PING
|
||||
|
||||
async def handle(self, conn, msg_json: Dict[str, Any]) -> None:
|
||||
"""
|
||||
处理PING消息,发送PONG响应
|
||||
消息格式:{"type": "ping"}
|
||||
Args:
|
||||
conn: WebSocket连接对象
|
||||
msg_json: PING消息的JSON数据
|
||||
"""
|
||||
# 检查是否启用了WebSocket心跳功能
|
||||
enable_websocket_ping = conn.config.get("enable_websocket_ping", False)
|
||||
if not enable_websocket_ping:
|
||||
conn.logger.debug(f"WebSocket心跳功能未启用,忽略PING消息")
|
||||
return
|
||||
|
||||
try:
|
||||
conn.logger.debug(f"收到PING消息,发送PONG响应")
|
||||
conn.last_activity_time = time.time() * 1000
|
||||
# 构造PONG响应消息
|
||||
pong_message = {
|
||||
"type": "pong",
|
||||
"timestamp": time.strftime("%Y-%m-%d %H:%M:%S", time.localtime()),
|
||||
}
|
||||
|
||||
# 发送PONG响应
|
||||
await conn.websocket.send(json.dumps(pong_message))
|
||||
|
||||
except Exception as e:
|
||||
conn.logger.error(f"处理PING消息时发生错误: {e}")
|
||||
@@ -7,6 +7,7 @@ from core.handle.textHandler.listenMessageHandler import ListenTextMessageHandle
|
||||
from core.handle.textHandler.mcpMessageHandler import McpTextMessageHandler
|
||||
from core.handle.textMessageHandler import TextMessageHandler
|
||||
from core.handle.textHandler.serverMessageHandler import ServerTextMessageHandler
|
||||
from core.handle.textHandler.pingMessageHandler import PingMessageHandler
|
||||
|
||||
TAG = __name__
|
||||
|
||||
@@ -27,6 +28,7 @@ class TextMessageHandlerRegistry:
|
||||
IotTextMessageHandler(),
|
||||
McpTextMessageHandler(),
|
||||
ServerTextMessageHandler(),
|
||||
PingMessageHandler(),
|
||||
]
|
||||
|
||||
for handler in handlers:
|
||||
|
||||
@@ -9,3 +9,4 @@ class TextMessageType(Enum):
|
||||
IOT = "iot"
|
||||
MCP = "mcp"
|
||||
SERVER = "server"
|
||||
PING = "ping"
|
||||
|
||||
@@ -33,38 +33,60 @@ class SimpleHttpServer:
|
||||
return f"ws://{local_ip}:{port}/xiaozhi/v1/"
|
||||
|
||||
async def start(self):
|
||||
server_config = self.config["server"]
|
||||
read_config_from_api = self.config.get("read_config_from_api", False)
|
||||
host = server_config.get("ip", "0.0.0.0")
|
||||
port = int(server_config.get("http_port", 8003))
|
||||
try:
|
||||
server_config = self.config["server"]
|
||||
read_config_from_api = self.config.get("read_config_from_api", False)
|
||||
host = server_config.get("ip", "0.0.0.0")
|
||||
port = int(server_config.get("http_port", 8003))
|
||||
|
||||
if port:
|
||||
app = web.Application()
|
||||
if port:
|
||||
app = web.Application()
|
||||
|
||||
if not read_config_from_api:
|
||||
# 如果没有开启智控台,只是单模块运行,就需要再添加简单OTA接口,用于下发websocket接口
|
||||
if not read_config_from_api:
|
||||
# 如果没有开启智控台,只是单模块运行,就需要再添加简单OTA接口,用于下发websocket接口
|
||||
app.add_routes(
|
||||
[
|
||||
web.get("/xiaozhi/ota/", self.ota_handler.handle_get),
|
||||
web.post("/xiaozhi/ota/", self.ota_handler.handle_post),
|
||||
web.options(
|
||||
"/xiaozhi/ota/", self.ota_handler.handle_options
|
||||
),
|
||||
# 下载接口,仅提供 data/bin/*.bin 下载
|
||||
web.get(
|
||||
"/xiaozhi/ota/download/{filename}",
|
||||
self.ota_handler.handle_download,
|
||||
),
|
||||
web.options(
|
||||
"/xiaozhi/ota/download/{filename}",
|
||||
self.ota_handler.handle_options,
|
||||
),
|
||||
]
|
||||
)
|
||||
# 添加路由
|
||||
app.add_routes(
|
||||
[
|
||||
web.get("/xiaozhi/ota/", self.ota_handler.handle_get),
|
||||
web.post("/xiaozhi/ota/", self.ota_handler.handle_post),
|
||||
web.options("/xiaozhi/ota/", self.ota_handler.handle_post),
|
||||
web.get("/mcp/vision/explain", self.vision_handler.handle_get),
|
||||
web.post(
|
||||
"/mcp/vision/explain", self.vision_handler.handle_post
|
||||
),
|
||||
web.options(
|
||||
"/mcp/vision/explain", self.vision_handler.handle_options
|
||||
),
|
||||
]
|
||||
)
|
||||
# 添加路由
|
||||
app.add_routes(
|
||||
[
|
||||
web.get("/mcp/vision/explain", self.vision_handler.handle_get),
|
||||
web.post("/mcp/vision/explain", self.vision_handler.handle_post),
|
||||
web.options("/mcp/vision/explain", self.vision_handler.handle_post),
|
||||
]
|
||||
)
|
||||
|
||||
# 运行服务
|
||||
runner = web.AppRunner(app)
|
||||
await runner.setup()
|
||||
site = web.TCPSite(runner, host, port)
|
||||
await site.start()
|
||||
# 运行服务
|
||||
runner = web.AppRunner(app)
|
||||
await runner.setup()
|
||||
site = web.TCPSite(runner, host, port)
|
||||
await site.start()
|
||||
|
||||
# 保持服务运行
|
||||
while True:
|
||||
await asyncio.sleep(3600) # 每隔 1 小时检查一次
|
||||
# 保持服务运行
|
||||
while True:
|
||||
await asyncio.sleep(3600) # 每隔 1 小时检查一次
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"HTTP服务器启动失败: {e}")
|
||||
import traceback
|
||||
|
||||
self.logger.bind(tag=TAG).error(f"错误堆栈: {traceback.format_exc()}")
|
||||
raise
|
||||
|
||||
@@ -0,0 +1,343 @@
|
||||
import json
|
||||
import uuid
|
||||
import asyncio
|
||||
import websockets
|
||||
import opuslib_next
|
||||
from typing import List
|
||||
from config.logger import setup_logging
|
||||
from core.providers.asr.base import ASRProviderBase
|
||||
from core.providers.asr.dto.dto import InterfaceType
|
||||
|
||||
TAG = __name__
|
||||
logger = setup_logging()
|
||||
|
||||
|
||||
class ASRProvider(ASRProviderBase):
|
||||
def __init__(self, config, delete_audio_file):
|
||||
super().__init__()
|
||||
self.interface_type = InterfaceType.STREAM
|
||||
self.config = config
|
||||
self.text = ""
|
||||
self.decoder = opuslib_next.Decoder(16000, 1)
|
||||
self.asr_ws = None
|
||||
self.forward_task = None
|
||||
self.is_processing = False
|
||||
self.server_ready = False # 服务器准备状态
|
||||
self.task_id = None # 当前任务ID
|
||||
|
||||
# 阿里百炼配置
|
||||
self.api_key = config.get("api_key")
|
||||
self.model = config.get("model", "paraformer-realtime-v2")
|
||||
self.sample_rate = config.get("sample_rate", 16000)
|
||||
self.format = config.get("format", "pcm")
|
||||
|
||||
# 可选参数
|
||||
self.vocabulary_id = config.get("vocabulary_id")
|
||||
self.disfluency_removal_enabled = config.get("disfluency_removal_enabled", False)
|
||||
self.language_hints = config.get("language_hints")
|
||||
self.semantic_punctuation_enabled = config.get("semantic_punctuation_enabled", False)
|
||||
max_sentence_silence = config.get("max_sentence_silence")
|
||||
self.max_sentence_silence = int(max_sentence_silence) if max_sentence_silence else 200
|
||||
self.multi_threshold_mode_enabled = config.get("multi_threshold_mode_enabled", False)
|
||||
self.punctuation_prediction_enabled = config.get("punctuation_prediction_enabled", True)
|
||||
self.inverse_text_normalization_enabled = config.get("inverse_text_normalization_enabled", True)
|
||||
|
||||
# WebSocket URL
|
||||
self.ws_url = "wss://dashscope.aliyuncs.com/api-ws/v1/inference"
|
||||
|
||||
self.output_dir = config.get("output_dir", "./audio_output")
|
||||
self.delete_audio_file = delete_audio_file
|
||||
|
||||
async def open_audio_channels(self, conn):
|
||||
await super().open_audio_channels(conn)
|
||||
|
||||
async def receive_audio(self, conn, audio, audio_have_voice):
|
||||
# 初始化音频缓存
|
||||
if not hasattr(conn, 'asr_audio_for_voiceprint'):
|
||||
conn.asr_audio_for_voiceprint = []
|
||||
|
||||
# 存储音频数据
|
||||
if audio:
|
||||
conn.asr_audio_for_voiceprint.append(audio)
|
||||
|
||||
conn.asr_audio.append(audio)
|
||||
conn.asr_audio = conn.asr_audio[-10:]
|
||||
|
||||
# 只在有声音且没有连接时建立连接
|
||||
if audio_have_voice and not self.is_processing and not self.asr_ws:
|
||||
try:
|
||||
await self._start_recognition(conn)
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"开始识别失败: {str(e)}")
|
||||
await self._cleanup()
|
||||
return
|
||||
|
||||
# 发送音频数据
|
||||
if self.asr_ws and self.is_processing and self.server_ready:
|
||||
try:
|
||||
pcm_frame = self.decoder.decode(audio, 960)
|
||||
# 直接发送PCM音频数据(二进制)
|
||||
await self.asr_ws.send(pcm_frame)
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).warning(f"发送音频失败: {str(e)}")
|
||||
await self._cleanup()
|
||||
|
||||
async def _start_recognition(self, conn):
|
||||
"""开始识别会话"""
|
||||
try:
|
||||
# 如果为手动模式,设置超时时长为最大值
|
||||
if conn.client_listen_mode == "manual":
|
||||
self.max_sentence_silence = 6000
|
||||
|
||||
self.is_processing = True
|
||||
self.task_id = uuid.uuid4().hex
|
||||
|
||||
# 建立WebSocket连接
|
||||
headers = {
|
||||
"Authorization": f"Bearer {self.api_key}"
|
||||
}
|
||||
|
||||
logger.bind(tag=TAG).debug(f"正在连接阿里百炼ASR服务, task_id: {self.task_id}")
|
||||
|
||||
self.asr_ws = await websockets.connect(
|
||||
self.ws_url,
|
||||
additional_headers=headers,
|
||||
max_size=1000000000,
|
||||
ping_interval=None,
|
||||
ping_timeout=None,
|
||||
close_timeout=5,
|
||||
)
|
||||
|
||||
logger.bind(tag=TAG).debug("WebSocket连接建立成功")
|
||||
|
||||
self.server_ready = False
|
||||
self.forward_task = asyncio.create_task(self._forward_results(conn))
|
||||
|
||||
# 发送run-task指令
|
||||
run_task_msg = self._build_run_task_message()
|
||||
await self.asr_ws.send(json.dumps(run_task_msg, ensure_ascii=False))
|
||||
logger.bind(tag=TAG).debug("已发送run-task指令,等待服务器准备...")
|
||||
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"建立ASR连接失败: {str(e)}")
|
||||
if self.asr_ws:
|
||||
await self.asr_ws.close()
|
||||
self.asr_ws = None
|
||||
self.is_processing = False
|
||||
raise
|
||||
|
||||
def _build_run_task_message(self) -> dict:
|
||||
"""构建run-task指令"""
|
||||
message = {
|
||||
"header": {
|
||||
"action": "run-task",
|
||||
"task_id": self.task_id,
|
||||
"streaming": "duplex"
|
||||
},
|
||||
"payload": {
|
||||
"task_group": "audio",
|
||||
"task": "asr",
|
||||
"function": "recognition",
|
||||
"model": self.model,
|
||||
"parameters": {
|
||||
"format": self.format,
|
||||
"sample_rate": self.sample_rate,
|
||||
"disfluency_removal_enabled": self.disfluency_removal_enabled,
|
||||
"semantic_punctuation_enabled": self.semantic_punctuation_enabled,
|
||||
"max_sentence_silence": self.max_sentence_silence,
|
||||
"multi_threshold_mode_enabled": self.multi_threshold_mode_enabled,
|
||||
"punctuation_prediction_enabled": self.punctuation_prediction_enabled,
|
||||
"inverse_text_normalization_enabled": self.inverse_text_normalization_enabled,
|
||||
},
|
||||
"input": {}
|
||||
}
|
||||
}
|
||||
|
||||
# 只有当模型名称以v2结尾时才添加vocabulary_id参数
|
||||
if self.model.lower().endswith("v2"):
|
||||
message["payload"]["parameters"]["vocabulary_id"] = self.vocabulary_id
|
||||
|
||||
if self.language_hints:
|
||||
message["payload"]["parameters"]["language_hints"] = self.language_hints
|
||||
|
||||
return message
|
||||
|
||||
async def _forward_results(self, conn):
|
||||
"""转发识别结果"""
|
||||
try:
|
||||
while not conn.stop_event.is_set():
|
||||
try:
|
||||
response = await asyncio.wait_for(self.asr_ws.recv(), timeout=1.0)
|
||||
result = json.loads(response)
|
||||
|
||||
header = result.get("header", {})
|
||||
payload = result.get("payload", {})
|
||||
event = header.get("event", "")
|
||||
|
||||
# 处理task-started事件
|
||||
if event == "task-started":
|
||||
self.server_ready = True
|
||||
logger.bind(tag=TAG).debug("服务器已准备,开始发送缓存音频...")
|
||||
|
||||
# 发送缓存音频
|
||||
if conn.asr_audio:
|
||||
for cached_audio in conn.asr_audio[-10:]:
|
||||
try:
|
||||
pcm_frame = self.decoder.decode(cached_audio, 960)
|
||||
await self.asr_ws.send(pcm_frame)
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).warning(f"发送缓存音频失败: {e}")
|
||||
break
|
||||
continue
|
||||
|
||||
# 处理result-generated事件
|
||||
elif event == "result-generated":
|
||||
output = payload.get("output", {})
|
||||
sentence = output.get("sentence", {})
|
||||
|
||||
text = sentence.get("text", "")
|
||||
sentence_end = sentence.get("sentence_end", False)
|
||||
end_time = sentence.get("end_time")
|
||||
|
||||
# 判断是否为最终结果(sentence_end为True且end_time不为null)
|
||||
is_final = sentence_end and end_time is not None
|
||||
|
||||
if is_final:
|
||||
logger.bind(tag=TAG).info(f"识别到文本: {text}")
|
||||
|
||||
# 手动模式下累积识别结果
|
||||
if conn.client_listen_mode == "manual":
|
||||
if self.text:
|
||||
self.text += text
|
||||
else:
|
||||
self.text = text
|
||||
|
||||
# 手动模式下,只有在收到stop信号后才触发处理
|
||||
if conn.client_voice_stop:
|
||||
audio_data = getattr(conn, 'asr_audio_for_voiceprint', [])
|
||||
if len(audio_data) > 0:
|
||||
logger.bind(tag=TAG).debug("收到最终识别结果,触发处理")
|
||||
await self.handle_voice_stop(conn, audio_data)
|
||||
# 清理音频缓存
|
||||
conn.asr_audio.clear()
|
||||
conn.reset_vad_states()
|
||||
break
|
||||
else:
|
||||
# 自动模式下直接覆盖
|
||||
self.text = text
|
||||
conn.reset_vad_states()
|
||||
audio_data = getattr(conn, 'asr_audio_for_voiceprint', [])
|
||||
await self.handle_voice_stop(conn, audio_data)
|
||||
break
|
||||
|
||||
# 处理task-finished事件
|
||||
elif event == "task-finished":
|
||||
logger.bind(tag=TAG).debug("任务已完成")
|
||||
break
|
||||
|
||||
# 处理task-failed事件
|
||||
elif event == "task-failed":
|
||||
error_code = header.get("error_code", "UNKNOWN")
|
||||
error_message = header.get("error_message", "未知错误")
|
||||
logger.bind(tag=TAG).error(f"任务失败: {error_code} - {error_message}")
|
||||
break
|
||||
|
||||
except asyncio.TimeoutError:
|
||||
continue
|
||||
except websockets.ConnectionClosed:
|
||||
logger.bind(tag=TAG).info("ASR服务连接已关闭")
|
||||
self.is_processing = False
|
||||
break
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"处理结果失败: {str(e)}")
|
||||
break
|
||||
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"结果转发失败: {str(e)}")
|
||||
finally:
|
||||
# 清理连接的音频缓存
|
||||
await self._cleanup()
|
||||
if conn:
|
||||
if hasattr(conn, 'asr_audio_for_voiceprint'):
|
||||
conn.asr_audio_for_voiceprint = []
|
||||
if hasattr(conn, 'asr_audio'):
|
||||
conn.asr_audio = []
|
||||
|
||||
async def _send_stop_request(self):
|
||||
"""发送停止请求(用于手动模式停止录音)"""
|
||||
if self.asr_ws:
|
||||
try:
|
||||
# 先停止音频发送
|
||||
self.is_processing = False
|
||||
|
||||
logger.bind(tag=TAG).debug("收到停止请求,发送finish-task指令")
|
||||
await self._send_finish_task()
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"发送停止请求失败: {e}")
|
||||
|
||||
async def _send_finish_task(self):
|
||||
"""发送finish-task指令"""
|
||||
if self.asr_ws and self.task_id:
|
||||
try:
|
||||
finish_msg = {
|
||||
"header": {
|
||||
"action": "finish-task",
|
||||
"task_id": self.task_id,
|
||||
"streaming": "duplex"
|
||||
},
|
||||
"payload": {
|
||||
"input": {}
|
||||
}
|
||||
}
|
||||
await self.asr_ws.send(json.dumps(finish_msg, ensure_ascii=False))
|
||||
logger.bind(tag=TAG).debug("已发送finish-task指令")
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"发送finish-task指令失败: {e}")
|
||||
|
||||
async def _cleanup(self):
|
||||
"""清理资源"""
|
||||
logger.bind(tag=TAG).debug(f"开始ASR会话清理 | 当前状态: processing={self.is_processing}, server_ready={self.server_ready}")
|
||||
|
||||
# 状态重置
|
||||
self.is_processing = False
|
||||
self.server_ready = False
|
||||
logger.bind(tag=TAG).debug("ASR状态已重置")
|
||||
|
||||
# 关闭连接
|
||||
if self.asr_ws:
|
||||
try:
|
||||
# 先发送finish-task指令
|
||||
await self._send_finish_task()
|
||||
# 等待一小段时间让服务器处理
|
||||
await asyncio.sleep(0.1)
|
||||
|
||||
logger.bind(tag=TAG).debug("正在关闭WebSocket连接")
|
||||
await asyncio.wait_for(self.asr_ws.close(), timeout=2.0)
|
||||
logger.bind(tag=TAG).debug("WebSocket连接已关闭")
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"关闭WebSocket连接失败: {e}")
|
||||
finally:
|
||||
self.asr_ws = None
|
||||
|
||||
# 清理任务引用
|
||||
self.forward_task = None
|
||||
self.task_id = None
|
||||
|
||||
logger.bind(tag=TAG).debug("ASR会话清理完成")
|
||||
|
||||
async def speech_to_text(self, opus_data, session_id, audio_format):
|
||||
"""获取识别结果"""
|
||||
result = self.text
|
||||
self.text = ""
|
||||
return result, None
|
||||
|
||||
async def close(self):
|
||||
"""关闭资源"""
|
||||
await self._cleanup()
|
||||
if hasattr(self, 'decoder') and self.decoder is not None:
|
||||
try:
|
||||
del self.decoder
|
||||
self.decoder = None
|
||||
logger.bind(tag=TAG).debug("Aliyun BL decoder resources released")
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).debug(f"释放Aliyun BL decoder资源时出错: {e}")
|
||||
@@ -33,7 +33,12 @@ class ASRProvider(ASRProviderBase):
|
||||
self.delete_audio_file = delete_audio_file
|
||||
|
||||
# 火山引擎ASR配置
|
||||
self.ws_url = "wss://openspeech.bytedance.com/api/v3/sauc/bigmodel"
|
||||
enable_multilingual = config.get("enable_multilingual", False)
|
||||
self.enable_multilingual = False if str(enable_multilingual).lower() == 'false' else True
|
||||
if self.enable_multilingual:
|
||||
self.ws_url = "wss://openspeech.bytedance.com/api/v3/sauc/bigmodel_nostream"
|
||||
else:
|
||||
self.ws_url = "wss://openspeech.bytedance.com/api/v3/sauc/bigmodel"
|
||||
self.uid = config.get("uid", "streaming_asr_service")
|
||||
self.workflow = config.get(
|
||||
"workflow", "audio_in,resample,partition,vad,fe,decode,itn,nlu_punctuate"
|
||||
@@ -42,11 +47,14 @@ class ASRProvider(ASRProviderBase):
|
||||
self.format = config.get("format", "pcm")
|
||||
self.codec = config.get("codec", "pcm")
|
||||
self.rate = config.get("sample_rate", 16000)
|
||||
self.language = config.get("language", "zh-CN")
|
||||
# language参数仅在多语种模式(bigmodel_nostream)下有效
|
||||
self.language = config.get("language") if self.enable_multilingual else None
|
||||
self.bits = config.get("bits", 16)
|
||||
self.channel = config.get("channel", 1)
|
||||
self.auth_method = config.get("auth_method", "token")
|
||||
self.secret = config.get("secret", "access_secret")
|
||||
end_window_size = config.get("end_window_size")
|
||||
self.end_window_size = int(end_window_size) if end_window_size else 200
|
||||
|
||||
async def open_audio_channels(self, conn):
|
||||
await super().open_audio_channels(conn)
|
||||
@@ -173,7 +181,8 @@ class ASRProvider(ASRProviderBase):
|
||||
utterances = payload["result"].get("utterances", [])
|
||||
# 检查duration和空文本的情况
|
||||
if (
|
||||
payload.get("audio_info", {}).get("duration", 0) > 2000
|
||||
not self.enable_multilingual # 注意:多语种模式不返回中间结果,需要等待最终结果
|
||||
and payload.get("audio_info", {}).get("duration", 0) > 2000
|
||||
and not utterances
|
||||
and not payload["result"].get("text")
|
||||
and conn.client_listen_mode != "manual"
|
||||
@@ -187,6 +196,10 @@ class ASRProvider(ASRProviderBase):
|
||||
|
||||
# 专门处理没有文本的识别结果(手动模式下可能已经识别完成但是没松按键)
|
||||
elif not payload["result"].get("text") and not utterances:
|
||||
# 多语种模式会持续返回空文本,直到最后返回完整结果,所以需要排除
|
||||
if self.enable_multilingual:
|
||||
continue
|
||||
|
||||
if conn.client_listen_mode == "manual" and conn.client_voice_stop and len(audio_data) > 0:
|
||||
logger.bind(tag=TAG).debug("消息结束收到停止信号,触发处理")
|
||||
await self.handle_voice_stop(conn, audio_data)
|
||||
@@ -291,18 +304,22 @@ class ASRProvider(ASRProviderBase):
|
||||
"sequence": 1,
|
||||
"boosting_table_name": self.boosting_table_name,
|
||||
"correct_table_name": self.correct_table_name,
|
||||
"end_window_size": 200,
|
||||
"end_window_size": self.end_window_size,
|
||||
},
|
||||
"audio": {
|
||||
"format": self.format,
|
||||
"codec": self.codec,
|
||||
"rate": self.rate,
|
||||
"language": self.language,
|
||||
"bits": self.bits,
|
||||
"channel": self.channel,
|
||||
"sample_rate": self.rate,
|
||||
},
|
||||
}
|
||||
|
||||
# language参数仅在多语种模式下添加
|
||||
if self.enable_multilingual and self.language:
|
||||
req["audio"]["language"] = self.language
|
||||
|
||||
logger.bind(tag=TAG).debug(
|
||||
f"构造请求参数: {json.dumps(req, ensure_ascii=False)}"
|
||||
)
|
||||
|
||||
@@ -14,7 +14,7 @@ class MemoryProviderBase(ABC):
|
||||
self.llm = llm
|
||||
|
||||
@abstractmethod
|
||||
async def save_memory(self, msgs):
|
||||
async def save_memory(self, msgs, session_id=None):
|
||||
"""Save a new memory for specific role and return memory ID"""
|
||||
print("this is base func", msgs)
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ class MemoryProvider(MemoryProviderBase):
|
||||
logger.bind(tag=TAG).error(f"详细错误: {traceback.format_exc()}")
|
||||
self.use_mem0 = False
|
||||
|
||||
async def save_memory(self, msgs):
|
||||
async def save_memory(self, msgs, session_id=None):
|
||||
if not self.use_mem0:
|
||||
return None
|
||||
if len(msgs) < 2:
|
||||
@@ -41,9 +41,7 @@ class MemoryProvider(MemoryProviderBase):
|
||||
for message in msgs
|
||||
if message.role != "system"
|
||||
]
|
||||
result = self.client.add(
|
||||
messages, user_id=self.role_id
|
||||
)
|
||||
result = self.client.add(messages, user_id=self.role_id)
|
||||
logger.bind(tag=TAG).debug(f"Save memory result: {result}")
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"保存记忆失败: {str(e)}")
|
||||
|
||||
@@ -4,7 +4,7 @@ import json
|
||||
import os
|
||||
import yaml
|
||||
from config.config_loader import get_project_dir
|
||||
from config.manage_api_client import save_mem_local_short
|
||||
from config.manage_api_client import generate_and_save_chat_summary
|
||||
import asyncio
|
||||
from core.utils.util import check_model_key
|
||||
|
||||
@@ -75,18 +75,6 @@ short_term_memory_prompt = """
|
||||
```
|
||||
"""
|
||||
|
||||
short_term_memory_prompt_only_content = """
|
||||
你是一个经验丰富的记忆总结者,擅长将对话内容进行总结摘要,遵循以下规则:
|
||||
1、总结user的重要信息,以便在未来的对话中提供更个性化的服务
|
||||
2、不要重复总结,不要遗忘之前记忆,除非原来的记忆超过了1800字内,否则不要遗忘、不要压缩用户的历史记忆
|
||||
3、用户操控的设备音量、播放音乐、天气、退出、不想对话等和用户本身无关的内容,这些信息不需要加入到总结中
|
||||
4、聊天内容中的今天的日期时间、今天的天气情况与用户事件无关的数据,这些信息如果当成记忆存储会影响后序对话,这些信息不需要加入到总结中
|
||||
5、不要把设备操控的成果结果和失败结果加入到总结中,也不要把用户的一些废话加入到总结中
|
||||
6、不要为了总结而总结,如果用户的聊天没有意义,请返回原来的历史记录也是可以的
|
||||
7、只需要返回总结摘要,严格控制在1800字内
|
||||
8、不要包含代码、xml,不需要解释、注释和说明,保存记忆时仅从对话提取信息,不要混入示例内容
|
||||
"""
|
||||
|
||||
|
||||
def extract_json_data(json_code):
|
||||
start = json_code.find("```json")
|
||||
@@ -144,7 +132,7 @@ class MemoryProvider(MemoryProviderBase):
|
||||
with open(self.memory_path, "w", encoding="utf-8") as f:
|
||||
yaml.dump(all_memory, f, allow_unicode=True)
|
||||
|
||||
async def save_memory(self, msgs):
|
||||
async def save_memory(self, msgs, session_id=None):
|
||||
# 打印使用的模型信息
|
||||
model_info = getattr(self.llm, "model_name", str(self.llm.__class__.__name__))
|
||||
logger.bind(tag=TAG).debug(f"使用记忆保存模型: {model_info}")
|
||||
@@ -188,20 +176,12 @@ class MemoryProvider(MemoryProviderBase):
|
||||
except Exception as e:
|
||||
print("Error:", e)
|
||||
else:
|
||||
result = self.llm.response_no_stream(
|
||||
short_term_memory_prompt_only_content,
|
||||
msgStr,
|
||||
max_tokens=2000,
|
||||
temperature=0.2,
|
||||
)
|
||||
# 使用异步版本,需要在事件循环中运行
|
||||
try:
|
||||
loop = asyncio.get_running_loop()
|
||||
loop.create_task(save_mem_local_short(self.role_id, result))
|
||||
except RuntimeError:
|
||||
# 如果没有运行中的事件循环,创建一个新的
|
||||
asyncio.run(save_mem_local_short(self.role_id, result))
|
||||
logger.bind(tag=TAG).info(f"Save memory successful - Role: {self.role_id}")
|
||||
# 当save_to_file为False时,调用Java端的聊天记录总结接口
|
||||
summary_id = session_id if session_id else self.role_id
|
||||
await generate_and_save_chat_summary(summary_id)
|
||||
logger.bind(tag=TAG).info(
|
||||
f"Save memory successful - Role: {self.role_id}, Session: {session_id}"
|
||||
)
|
||||
|
||||
return self.short_memory
|
||||
|
||||
|
||||
@@ -11,7 +11,7 @@ class MemoryProvider(MemoryProviderBase):
|
||||
def __init__(self, config, summary_memory=None):
|
||||
super().__init__(config)
|
||||
|
||||
async def save_memory(self, msgs):
|
||||
async def save_memory(self, msgs, session_id=None):
|
||||
logger.bind(tag=TAG).debug("nomem mode: No memory saving is performed.")
|
||||
return None
|
||||
|
||||
|
||||
@@ -3,12 +3,8 @@
|
||||
import asyncio
|
||||
import os
|
||||
import json
|
||||
from datetime import timedelta
|
||||
from typing import Dict, Any, List
|
||||
|
||||
from mcp import Implementation
|
||||
from mcp.client.session import SamplingFnT, ElicitationFnT, ListRootsFnT, LoggingFnT, MessageHandlerFnT
|
||||
from mcp.shared.session import ProgressFnT
|
||||
from mcp.types import LoggingMessageNotificationParams
|
||||
|
||||
from config.config_loader import get_project_dir
|
||||
@@ -33,6 +29,7 @@ class ServerMCPManager:
|
||||
)
|
||||
self.clients: Dict[str, ServerMCPClient] = {}
|
||||
self.tools = []
|
||||
self._init_lock = asyncio.Lock()
|
||||
|
||||
def load_config(self) -> Dict[str, Any]:
|
||||
"""加载MCP服务配置"""
|
||||
@@ -49,29 +46,50 @@ class ServerMCPManager:
|
||||
)
|
||||
return {}
|
||||
|
||||
async def _init_server(self, name: str, srv_config: Dict[str, Any]):
|
||||
"""初始化单个MCP服务"""
|
||||
client = None
|
||||
try:
|
||||
# 初始化服务端MCP客户端
|
||||
logger.bind(tag=TAG).info(f"初始化服务端MCP客户端: {name}")
|
||||
client = ServerMCPClient(srv_config)
|
||||
# 设置超时时间10秒
|
||||
await asyncio.wait_for(client.initialize(logging_callback=self.logging_callback), timeout=10)
|
||||
|
||||
# 使用锁保护共享状态的修改
|
||||
async with self._init_lock:
|
||||
self.clients[name] = client
|
||||
client_tools = client.get_available_tools()
|
||||
self.tools.extend(client_tools)
|
||||
|
||||
except asyncio.TimeoutError:
|
||||
logger.bind(tag=TAG).error(
|
||||
f"Failed to initialize MCP server {name}: Timeout"
|
||||
)
|
||||
if client:
|
||||
await client.cleanup()
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(
|
||||
f"Failed to initialize MCP server {name}: {e}"
|
||||
)
|
||||
if client:
|
||||
await client.cleanup()
|
||||
|
||||
async def initialize_servers(self) -> None:
|
||||
"""初始化所有MCP服务"""
|
||||
config = self.load_config()
|
||||
tasks = []
|
||||
for name, srv_config in config.items():
|
||||
if not srv_config.get("command") and not srv_config.get("url"):
|
||||
logger.bind(tag=TAG).warning(
|
||||
f"Skipping server {name}: neither command nor url specified"
|
||||
)
|
||||
continue
|
||||
|
||||
try:
|
||||
# 初始化服务端MCP客户端
|
||||
logger.bind(tag=TAG).info(f"初始化服务端MCP客户端: {name}")
|
||||
client = ServerMCPClient(srv_config)
|
||||
await client.initialize(logging_callback=self.logging_callback)
|
||||
self.clients[name] = client
|
||||
client_tools = client.get_available_tools()
|
||||
self.tools.extend(client_tools)
|
||||
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(
|
||||
f"Failed to initialize MCP server {name}: {e}"
|
||||
)
|
||||
|
||||
tasks.append(self._init_server(name, srv_config))
|
||||
|
||||
if tasks:
|
||||
await asyncio.gather(*tasks)
|
||||
|
||||
# 输出当前支持的服务端MCP工具列表
|
||||
if hasattr(self.conn, "func_handler") and self.conn.func_handler:
|
||||
|
||||
@@ -160,10 +160,16 @@ class TTSProvider(TTSProviderBase):
|
||||
self.speech_rate = int(speech_rate) if speech_rate else 0
|
||||
self.loudness_rate = int(loudness_rate) if loudness_rate else 0
|
||||
self.pitch = int(pitch) if pitch else 0
|
||||
# 多情感音色参数
|
||||
self.emotion = config.get("emotion", "neutral")
|
||||
emotion_scale = config.get("emotion_scale", "4")
|
||||
self.emotion_scale = int(emotion_scale) if emotion_scale else 4
|
||||
|
||||
self.ws_url = config.get("ws_url")
|
||||
self.authorization = config.get("authorization")
|
||||
self.header = {"Authorization": f"{self.authorization}{self.access_token}"}
|
||||
self.enable_two_way = True
|
||||
enable_ws_reuse_value = config.get("enable_ws_reuse", True)
|
||||
self.enable_ws_reuse = False if str(enable_ws_reuse_value).lower() == 'false' else True
|
||||
self.tts_text = ""
|
||||
self.opus_encoder = opus_encoder_utils.OpusEncoderUtils(
|
||||
sample_rate=16000, channels=1, frame_size_ms=60
|
||||
@@ -184,8 +190,14 @@ class TTSProvider(TTSProviderBase):
|
||||
"""建立新的WebSocket连接,并启动监听任务(仅第一次)"""
|
||||
try:
|
||||
if self.ws:
|
||||
logger.bind(tag=TAG).info(f"使用已有链接...")
|
||||
return self.ws
|
||||
if self.enable_ws_reuse:
|
||||
logger.bind(tag=TAG).info(f"使用已有链接...")
|
||||
return self.ws
|
||||
else:
|
||||
try:
|
||||
await self.finish_connection()
|
||||
except:
|
||||
pass
|
||||
logger.bind(tag=TAG).debug("开始建立新连接...")
|
||||
ws_header = {
|
||||
"X-Api-App-Key": self.appId,
|
||||
@@ -208,6 +220,22 @@ class TTSProvider(TTSProviderBase):
|
||||
logger.bind(tag=TAG).error(f"建立连接失败: {str(e)}")
|
||||
self.ws = None
|
||||
raise
|
||||
|
||||
async def finish_connection(self):
|
||||
"""发送 FinishConnection 事件,等待服务端返回 EVENT_ConnectionFinished"""
|
||||
try:
|
||||
if self.ws:
|
||||
logger.bind(tag=TAG).debug("开始关闭连接...")
|
||||
header = Header(
|
||||
message_type=FULL_CLIENT_REQUEST,
|
||||
message_type_specific_flags=MsgTypeFlagWithEvent,
|
||||
serial_method=JSON,
|
||||
).as_bytes()
|
||||
optional = Optional(event=EVENT_FinishConnection).as_bytes()
|
||||
payload = str.encode("{}")
|
||||
await self.send_event(self.ws, header, optional, payload)
|
||||
except:
|
||||
pass
|
||||
|
||||
def tts_text_priority_thread(self):
|
||||
"""火山引擎双流式TTS的文本处理线程"""
|
||||
@@ -224,10 +252,16 @@ class TTSProvider(TTSProviderBase):
|
||||
if self.conn.client_abort:
|
||||
try:
|
||||
logger.bind(tag=TAG).info("收到打断信息,终止TTS文本处理线程")
|
||||
asyncio.run_coroutine_threadsafe(
|
||||
self.cancel_session(self.conn.sentence_id),
|
||||
loop=self.conn.loop,
|
||||
)
|
||||
if self.enable_ws_reuse:
|
||||
asyncio.run_coroutine_threadsafe(
|
||||
self.cancel_session(self.conn.sentence_id),
|
||||
loop=self.conn.loop,
|
||||
)
|
||||
else:
|
||||
asyncio.run_coroutine_threadsafe(
|
||||
self.finish_connection(),
|
||||
loop=self.conn.loop,
|
||||
)
|
||||
continue
|
||||
except Exception as e:
|
||||
logger.bind(tag=TAG).error(f"取消TTS会话失败: {str(e)}")
|
||||
@@ -432,6 +466,11 @@ class TTSProvider(TTSProviderBase):
|
||||
res = self.parser_response(msg)
|
||||
self.print_response(res, "send_text res:")
|
||||
|
||||
# 优先处理连接级别事件
|
||||
if res.optional.event == EVENT_ConnectionFinished:
|
||||
logger.bind(tag=TAG).debug(f"链接关闭成功~~")
|
||||
break
|
||||
|
||||
# 只处理当前活跃会话的响应
|
||||
if res.optional.sessionId and self.conn.sentence_id != res.optional.sessionId:
|
||||
# 如果是会话结束相关事件,即使会话ID不匹配也要重置状态
|
||||
@@ -461,6 +500,9 @@ class TTSProvider(TTSProviderBase):
|
||||
logger.bind(tag=TAG).debug(f"会话结束~~")
|
||||
self.activate_session = False
|
||||
self._process_before_stop_play_files()
|
||||
# 非复用模式下,会话结束后发送 FinishConnection
|
||||
if not self.enable_ws_reuse:
|
||||
await self.finish_connection()
|
||||
except websockets.ConnectionClosed:
|
||||
logger.bind(tag=TAG).warning("WebSocket连接已关闭")
|
||||
break
|
||||
@@ -479,6 +521,7 @@ class TTSProvider(TTSProviderBase):
|
||||
self.ws = None
|
||||
# 监听任务退出时清理引用
|
||||
finally:
|
||||
self.activate_session = False
|
||||
self._monitor_task = None
|
||||
|
||||
async def send_event(
|
||||
@@ -604,6 +647,19 @@ class TTSProvider(TTSProviderBase):
|
||||
audio_format="pcm",
|
||||
audio_sample_rate=16000,
|
||||
):
|
||||
audio_params = {
|
||||
"format": audio_format,
|
||||
"sample_rate": audio_sample_rate,
|
||||
"speech_rate": self.speech_rate,
|
||||
"loudness_rate": self.loudness_rate
|
||||
}
|
||||
|
||||
# 如果是多情感音色,添加情感参数
|
||||
if '_emo_' in self.voice:
|
||||
if self.emotion:
|
||||
audio_params["emotion"] = self.emotion
|
||||
audio_params["emotion_scale"] = self.emotion_scale
|
||||
|
||||
return str.encode(
|
||||
json.dumps(
|
||||
{
|
||||
@@ -613,12 +669,7 @@ class TTSProvider(TTSProviderBase):
|
||||
"req_params": {
|
||||
"text": text,
|
||||
"speaker": speaker,
|
||||
"audio_params": {
|
||||
"format": audio_format,
|
||||
"sample_rate": audio_sample_rate,
|
||||
"speech_rate": self.speech_rate,
|
||||
"loudness_rate": self.loudness_rate
|
||||
},
|
||||
"audio_params": audio_params,
|
||||
"additions": json.dumps({
|
||||
"post_process": {
|
||||
"pitch": self.pitch
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import time
|
||||
import asyncio
|
||||
from collections import deque
|
||||
from config.logger import setup_logging
|
||||
|
||||
TAG = __name__
|
||||
@@ -18,13 +19,14 @@ class AudioRateController:
|
||||
frame_duration: 单个音频帧时长(毫秒),默认60ms
|
||||
"""
|
||||
self.frame_duration = frame_duration
|
||||
self.queue = []
|
||||
self.queue = deque()
|
||||
self.play_position = 0 # 虚拟播放位置(毫秒)
|
||||
self.start_timestamp = None # 开始时间戳(只读,不修改)
|
||||
self.pending_send_task = None
|
||||
self.logger = logger
|
||||
self.queue_empty_event = asyncio.Event() # 队列清空事件
|
||||
self.queue_empty_event.set() # 初始为空状态
|
||||
self.queue_has_data_event = asyncio.Event() # 队列数据事件
|
||||
|
||||
def reset(self):
|
||||
"""重置控制器状态"""
|
||||
@@ -34,13 +36,17 @@ class AudioRateController:
|
||||
|
||||
self.queue.clear()
|
||||
self.play_position = 0
|
||||
self.start_timestamp = time.time()
|
||||
self.queue_empty_event.set() # 队列已清空
|
||||
self.start_timestamp = None # 由首个音频包设置
|
||||
# 相关事件处理
|
||||
self.queue_empty_event.set()
|
||||
self.queue_has_data_event.clear()
|
||||
|
||||
def add_audio(self, opus_packet):
|
||||
"""添加音频包到队列"""
|
||||
self.queue.append(("audio", opus_packet))
|
||||
self.queue_empty_event.clear() # 队列非空,清除事件
|
||||
# 相关事件处理
|
||||
self.queue_empty_event.clear()
|
||||
self.queue_has_data_event.set()
|
||||
|
||||
def add_message(self, message_callback):
|
||||
"""
|
||||
@@ -50,13 +56,15 @@ class AudioRateController:
|
||||
message_callback: 消息发送回调函数 async def()
|
||||
"""
|
||||
self.queue.append(("message", message_callback))
|
||||
self.queue_empty_event.clear() # 队列非空,清除事件
|
||||
# 相关事件处理
|
||||
self.queue_empty_event.clear()
|
||||
self.queue_has_data_event.set()
|
||||
|
||||
def _get_elapsed_ms(self):
|
||||
"""获取已经过的时间(毫秒)"""
|
||||
if self.start_timestamp is None:
|
||||
return 0
|
||||
return (time.time() - self.start_timestamp) * 1000
|
||||
return (time.monotonic() - self.start_timestamp) * 1000
|
||||
|
||||
async def check_queue(self, send_audio_callback):
|
||||
"""
|
||||
@@ -65,9 +73,6 @@ class AudioRateController:
|
||||
Args:
|
||||
send_audio_callback: 发送音频的回调函数 async def(opus_packet)
|
||||
"""
|
||||
if self.start_timestamp is None:
|
||||
self.start_timestamp = time.time()
|
||||
|
||||
while self.queue:
|
||||
item = self.queue[0]
|
||||
item_type = item[0]
|
||||
@@ -75,7 +80,7 @@ class AudioRateController:
|
||||
if item_type == "message":
|
||||
# 消息类型:立即发送,不占用播放时间
|
||||
_, message_callback = item
|
||||
self.queue.pop(0)
|
||||
self.queue.popleft()
|
||||
try:
|
||||
await message_callback()
|
||||
except Exception as e:
|
||||
@@ -83,6 +88,9 @@ class AudioRateController:
|
||||
raise
|
||||
|
||||
elif item_type == "audio":
|
||||
if self.start_timestamp is None:
|
||||
self.start_timestamp = time.monotonic()
|
||||
|
||||
_, opus_packet = item
|
||||
|
||||
# 循环等待直到时间到达
|
||||
@@ -107,16 +115,17 @@ class AudioRateController:
|
||||
break
|
||||
|
||||
# 时间已到,从队列移除并发送
|
||||
self.queue.pop(0)
|
||||
self.queue.popleft()
|
||||
self.play_position += self.frame_duration
|
||||
|
||||
try:
|
||||
await send_audio_callback(opus_packet)
|
||||
except Exception as e:
|
||||
self.logger.bind(tag=TAG).error(f"发送音频失败: {e}")
|
||||
raise
|
||||
|
||||
# 队列处理完后清除事件
|
||||
self.queue_empty_event.set()
|
||||
self.queue_has_data_event.clear()
|
||||
|
||||
def start_sending(self, send_audio_callback):
|
||||
"""
|
||||
@@ -132,9 +141,10 @@ class AudioRateController:
|
||||
async def _send_loop():
|
||||
try:
|
||||
while True:
|
||||
# 等待队列数据事件,不轮询等待占用CPU
|
||||
await self.queue_has_data_event.wait()
|
||||
|
||||
await self.check_queue(send_audio_callback)
|
||||
# 如果队列空了,短暂等待后再检查(避免 busy loop)
|
||||
await asyncio.sleep(0.01)
|
||||
except asyncio.CancelledError:
|
||||
self.logger.bind(tag=TAG).debug("音频发送循环已停止")
|
||||
except Exception as e:
|
||||
|
||||
@@ -184,10 +184,25 @@ class PromptManager:
|
||||
def update_context_info(self, conn, client_ip: str):
|
||||
"""同步更新上下文信息"""
|
||||
try:
|
||||
# 获取位置信息(使用全局缓存)
|
||||
local_address = self._get_location_info(client_ip)
|
||||
# 获取天气信息(使用全局缓存)
|
||||
self._get_weather_info(conn, local_address)
|
||||
local_address = ""
|
||||
if (
|
||||
client_ip
|
||||
and self.base_prompt_template
|
||||
and (
|
||||
"local_address" in self.base_prompt_template
|
||||
or "weather_info" in self.base_prompt_template
|
||||
)
|
||||
):
|
||||
# 获取位置信息(使用全局缓存)
|
||||
local_address = self._get_location_info(client_ip)
|
||||
|
||||
if (
|
||||
self.base_prompt_template
|
||||
and "weather_info" in self.base_prompt_template
|
||||
and local_address
|
||||
):
|
||||
# 获取天气信息(使用全局缓存)
|
||||
self._get_weather_info(conn, local_address)
|
||||
|
||||
# 获取配置的上下文数据
|
||||
if hasattr(conn, "device_id") and conn.device_id:
|
||||
|
||||
@@ -195,7 +195,7 @@ def get_news_from_chinanews(
|
||||
|
||||
# 否则,获取新闻列表并随机选择一条
|
||||
# 从配置中获取RSS URL
|
||||
rss_config = conn.config["plugins"]["get_news_from_chinanews"]
|
||||
rss_config = conn.config.get("plugins", {}).get("get_news_from_chinanews", {})
|
||||
default_rss_url = rss_config.get(
|
||||
"default_rss_url", "https://www.chinanews.com.cn/rss/society.xml"
|
||||
)
|
||||
|
||||
@@ -120,10 +120,10 @@ def fetch_news_from_api(conn, source="thepaper"):
|
||||
"""从API获取新闻列表"""
|
||||
try:
|
||||
api_url = f"https://newsnow.busiyi.world/api/s?id={source}"
|
||||
if conn.config["plugins"].get("get_news_from_newsnow") and conn.config[
|
||||
"plugins"
|
||||
]["get_news_from_newsnow"].get("url"):
|
||||
api_url = conn.config["plugins"]["get_news_from_newsnow"]["url"] + source
|
||||
|
||||
news_config = conn.config.get("plugins", {}).get("get_news_from_newsnow", {})
|
||||
if news_config.get("url"):
|
||||
api_url = news_config["url"] + source
|
||||
|
||||
headers = {"User-Agent": "Mozilla/5.0"}
|
||||
response = requests.get(api_url, headers=headers, timeout=10)
|
||||
|
||||
@@ -158,13 +158,10 @@ def parse_weather_info(soup):
|
||||
def get_weather(conn, location: str = None, lang: str = "zh_CN"):
|
||||
from core.utils.cache.manager import cache_manager, CacheType
|
||||
|
||||
api_host = conn.config["plugins"]["get_weather"].get(
|
||||
"api_host", "mj7p3y7naa.re.qweatherapi.com"
|
||||
)
|
||||
api_key = conn.config["plugins"]["get_weather"].get(
|
||||
"api_key", "a861d0d5e7bf4ee1a83d9a9e4f96d4da"
|
||||
)
|
||||
default_location = conn.config["plugins"]["get_weather"]["default_location"]
|
||||
weather_config = conn.config.get("plugins", {}).get("get_weather", {})
|
||||
api_host = weather_config.get("api_host", "mj7p3y7naa.re.qweatherapi.com")
|
||||
api_key = weather_config.get("api_key", "a861d0d5e7bf4ee1a83d9a9e4f96d4da")
|
||||
default_location = weather_config.get("default_location", "广州")
|
||||
client_ip = conn.client_ip
|
||||
|
||||
# 优先使用用户提供的location参数
|
||||
|
||||
@@ -11,15 +11,17 @@ def append_devices_to_prompt(conn):
|
||||
"functions", []
|
||||
)
|
||||
|
||||
# 安全地获取插件配置
|
||||
plugins_config = conn.config.get("plugins", {})
|
||||
config_source = (
|
||||
"home_assistant"
|
||||
if conn.config["plugins"].get("home_assistant")
|
||||
if plugins_config.get("home_assistant")
|
||||
else "hass_get_state"
|
||||
)
|
||||
|
||||
if "hass_get_state" in funcs or "hass_set_state" in funcs:
|
||||
prompt = "\n下面是我家智能设备列表(位置,设备名,entity_id),可以通过homeassistant控制\n"
|
||||
deviceStr = conn.config["plugins"].get(config_source, {}).get("devices", "")
|
||||
deviceStr = plugins_config.get(config_source, {}).get("devices", "")
|
||||
conn.prompt += prompt + deviceStr + "\n"
|
||||
# 更新提示词
|
||||
conn.dialogue.update_system_message(conn.prompt)
|
||||
@@ -30,17 +32,17 @@ def initialize_hass_handler(conn):
|
||||
if not conn.load_function_plugin:
|
||||
return ha_config
|
||||
|
||||
# 安全地获取插件配置
|
||||
plugins_config = conn.config.get("plugins", {})
|
||||
# 确定配置来源
|
||||
config_source = (
|
||||
"home_assistant"
|
||||
if conn.config["plugins"].get("home_assistant")
|
||||
else "hass_get_state"
|
||||
"home_assistant" if plugins_config.get("home_assistant") else "hass_get_state"
|
||||
)
|
||||
if not conn.config["plugins"].get(config_source):
|
||||
if not plugins_config.get(config_source):
|
||||
return ha_config
|
||||
|
||||
# 统一获取配置
|
||||
plugin_config = conn.config["plugins"][config_source]
|
||||
plugin_config = plugins_config[config_source]
|
||||
ha_config["base_url"] = plugin_config.get("base_url")
|
||||
ha_config["api_key"] = plugin_config.get("api_key")
|
||||
|
||||
|
||||
@@ -118,8 +118,9 @@ def get_music_files(music_dir, music_ext):
|
||||
def initialize_music_handler(conn):
|
||||
global MUSIC_CACHE
|
||||
if MUSIC_CACHE == {}:
|
||||
if "play_music" in conn.config["plugins"]:
|
||||
MUSIC_CACHE["music_config"] = conn.config["plugins"]["play_music"]
|
||||
plugins_config = conn.config.get("plugins", {})
|
||||
if "play_music" in plugins_config:
|
||||
MUSIC_CACHE["music_config"] = plugins_config["play_music"]
|
||||
MUSIC_CACHE["music_dir"] = os.path.abspath(
|
||||
MUSIC_CACHE["music_config"].get("music_dir", "./music") # 默认路径修改
|
||||
)
|
||||
|
||||
@@ -32,9 +32,10 @@ def search_from_ragflow(conn, question=None):
|
||||
else:
|
||||
question = str(question) if question is not None else ""
|
||||
|
||||
base_url = conn.config["plugins"]["search_from_ragflow"].get("base_url", "")
|
||||
api_key = conn.config["plugins"]["search_from_ragflow"].get("api_key", "")
|
||||
dataset_ids = conn.config["plugins"]["search_from_ragflow"].get("dataset_ids", [])
|
||||
ragflow_config = conn.config.get("plugins", {}).get("search_from_ragflow", {})
|
||||
base_url = ragflow_config.get("base_url", "")
|
||||
api_key = ragflow_config.get("api_key", "")
|
||||
dataset_ids = ragflow_config.get("dataset_ids", [])
|
||||
|
||||
url = base_url + "/api/v1/retrieval"
|
||||
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
|
||||
@@ -64,12 +65,24 @@ def search_from_ragflow(conn, question=None):
|
||||
result = json.loads(response_text)
|
||||
|
||||
if result.get("code") != 0:
|
||||
error_detail = response.get("error", {}).get("detail", "")
|
||||
error_detail = result.get("error", {}).get("detail", "未知错误")
|
||||
error_message = result.get("error", {}).get("message", "")
|
||||
error_code = result.get("code", "")
|
||||
|
||||
# 安全地记录错误信息
|
||||
logger.bind(tag=TAG).error(
|
||||
"从RAGflow获取信息失败,原因:%s", str(error_detail)
|
||||
f"RAGFlow API调用失败,响应码:{error_code},错误详情:{error_detail},完整响应:{result}"
|
||||
)
|
||||
return ActionResponse(Action.RESPONSE, None, "RAG接口返回异常")
|
||||
|
||||
# 构建详细的错误响应
|
||||
error_response = f"RAG接口返回异常(错误码:{error_code})"
|
||||
|
||||
if error_message:
|
||||
error_response += f":{error_message}"
|
||||
if error_detail:
|
||||
error_response += f"\n详情:{error_detail}"
|
||||
|
||||
return ActionResponse(Action.RESPONSE, None, error_response)
|
||||
|
||||
chunks = result.get("data", {}).get("chunks", [])
|
||||
contents = []
|
||||
@@ -93,7 +106,57 @@ def search_from_ragflow(conn, question=None):
|
||||
context_text = "根据知识库查询结果,没有相关信息。"
|
||||
return ActionResponse(Action.REQLLM, context_text, None)
|
||||
|
||||
except requests.exceptions.RequestException as e:
|
||||
# 网络请求异常
|
||||
error_type = type(e).__name__
|
||||
logger.bind(tag=TAG).error(
|
||||
f"RAGflow网络请求失败,异常类型:{error_type},详情:{str(e)}"
|
||||
)
|
||||
|
||||
# 根据异常类型提供更详细的错误信息和解决方案
|
||||
if isinstance(e, requests.exceptions.ConnectTimeout):
|
||||
error_response = "RAG接口连接超时(5秒)"
|
||||
error_response += "\n可能原因:RAGflow服务未启动或网络连接问题"
|
||||
error_response += "\n解决方案:请检查RAGflow服务状态和网络连接"
|
||||
|
||||
elif isinstance(e, requests.exceptions.ConnectionError):
|
||||
error_response = "无法连接到RAG接口"
|
||||
error_response += "\n可能原因:RAGflow服务地址错误或服务未运行"
|
||||
error_response += "\n解决方案:请检查RAGflow服务地址配置和服务状态"
|
||||
|
||||
elif isinstance(e, requests.exceptions.Timeout):
|
||||
error_response = "RAG接口请求超时"
|
||||
error_response += "\n可能原因:RAGflow服务响应缓慢或网络延迟"
|
||||
error_response += "\n解决方案:请稍后重试或检查RAGflow服务性能"
|
||||
|
||||
elif isinstance(e, requests.exceptions.HTTPError):
|
||||
# 处理HTTP错误状态码
|
||||
if hasattr(e.response, "status_code"):
|
||||
status_code = e.response.status_code
|
||||
error_response = f"RAG接口HTTP错误(状态码:{status_code})"
|
||||
|
||||
# 尝试获取响应内容中的错误信息
|
||||
try:
|
||||
error_detail = e.response.json().get("error", {}).get("message", "")
|
||||
if error_detail:
|
||||
error_response += f"\n错误详情:{error_detail}"
|
||||
except:
|
||||
pass
|
||||
else:
|
||||
error_response = f"RAG接口HTTP异常:{str(e)}"
|
||||
|
||||
else:
|
||||
error_response = f"RAG接口网络异常({error_type}):{str(e)}"
|
||||
|
||||
return ActionResponse(Action.RESPONSE, None, error_response)
|
||||
|
||||
except Exception as e:
|
||||
# 使用安全的方式记录异常,避免编码问题
|
||||
logger.bind(tag=TAG).error("从RAGflow获取信息失败,原因:%s", str(e))
|
||||
return ActionResponse(Action.RESPONSE, None, "RAG接口返回异常")
|
||||
# 其他异常
|
||||
error_type = type(e).__name__
|
||||
logger.bind(tag=TAG).error(
|
||||
f"RAGflow处理异常,异常类型:{error_type},详情:{str(e)}"
|
||||
)
|
||||
|
||||
# 提供详细的错误信息
|
||||
error_response = f"RAG接口处理异常({error_type}):{str(e)}"
|
||||
return ActionResponse(Action.RESPONSE, None, error_response)
|
||||
|
||||
@@ -24,9 +24,9 @@ requests==2.32.5
|
||||
cozepy==0.20.0
|
||||
mem0ai==1.0.0
|
||||
bs4==0.0.2
|
||||
modelscope==1.23.2
|
||||
modelscope==1.32.0
|
||||
sherpa_onnx==1.12.17
|
||||
mcp==1.20.0
|
||||
mcp==1.22.0
|
||||
cnlunar==0.2.0
|
||||
PySocks==1.7.1
|
||||
dashscope==1.25.2
|
||||
@@ -36,7 +36,7 @@ aioconsole==0.8.2
|
||||
markitdown==0.1.3
|
||||
mcp-proxy==0.10.0
|
||||
PyJWT==2.10.1
|
||||
psutil==7.0.0
|
||||
psutil==7.1.3
|
||||
portalocker==3.2.0
|
||||
Jinja2==3.1.6
|
||||
vosk==0.3.45
|
||||
@@ -267,6 +267,7 @@ span.connection-status.llm-emoji {
|
||||
.llm-emoji .status {
|
||||
font-size: 14px !important;
|
||||
padding: 8px 20px !important;
|
||||
line-height: 1.2 !important;
|
||||
}
|
||||
|
||||
.emoji-large {
|
||||
@@ -393,7 +394,7 @@ span.connection-status.llm-emoji {
|
||||
/* ==================== 会话记录和日志 ==================== */
|
||||
.flex-container {
|
||||
display: flex;
|
||||
margin-top: 10px;
|
||||
margin-top: 20px;
|
||||
background-color: #f9fafb;
|
||||
}
|
||||
|
||||
|
||||
@@ -18,7 +18,6 @@ export function loadConfig() {
|
||||
const deviceMacInput = document.getElementById('deviceMac');
|
||||
const deviceNameInput = document.getElementById('deviceName');
|
||||
const clientIdInput = document.getElementById('clientId');
|
||||
const tokenInput = document.getElementById('token');
|
||||
const otaUrlInput = document.getElementById('otaUrl');
|
||||
|
||||
// 从localStorage加载MAC地址,如果没有则生成新的
|
||||
@@ -40,11 +39,6 @@ export function loadConfig() {
|
||||
clientIdInput.value = savedClientId;
|
||||
}
|
||||
|
||||
const savedToken = localStorage.getItem('xz_tester_token');
|
||||
if (savedToken) {
|
||||
tokenInput.value = savedToken;
|
||||
}
|
||||
|
||||
const savedOtaUrl = localStorage.getItem('xz_tester_otaUrl');
|
||||
if (savedOtaUrl) {
|
||||
otaUrlInput.value = savedOtaUrl;
|
||||
@@ -56,12 +50,10 @@ export function saveConfig() {
|
||||
const deviceMacInput = document.getElementById('deviceMac');
|
||||
const deviceNameInput = document.getElementById('deviceName');
|
||||
const clientIdInput = document.getElementById('clientId');
|
||||
const tokenInput = document.getElementById('token');
|
||||
|
||||
localStorage.setItem('xz_tester_deviceMac', deviceMacInput.value);
|
||||
localStorage.setItem('xz_tester_deviceName', deviceNameInput.value);
|
||||
localStorage.setItem('xz_tester_clientId', clientIdInput.value);
|
||||
localStorage.setItem('xz_tester_token', tokenInput.value);
|
||||
}
|
||||
|
||||
// 获取配置值
|
||||
@@ -71,8 +63,7 @@ export function getConfig() {
|
||||
deviceId: deviceMac, // 使用MAC地址作为deviceId
|
||||
deviceName: document.getElementById('deviceName').value.trim(),
|
||||
deviceMac: deviceMac,
|
||||
clientId: document.getElementById('clientId').value.trim(),
|
||||
token: document.getElementById('token').value.trim()
|
||||
clientId: document.getElementById('clientId').value.trim()
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -249,6 +249,41 @@ export class AudioPlayer {
|
||||
this.playBufferedAudio();
|
||||
this.startAudioBuffering();
|
||||
}
|
||||
|
||||
// 获取音频包统计信息
|
||||
getAudioStats() {
|
||||
if (!this.streamingContext) {
|
||||
return {
|
||||
pendingDecode: 0,
|
||||
pendingPlay: 0,
|
||||
totalPending: 0
|
||||
};
|
||||
}
|
||||
|
||||
const pendingDecode = this.streamingContext.getPendingDecodeCount();
|
||||
const pendingPlay = this.streamingContext.getPendingPlayCount();
|
||||
|
||||
return {
|
||||
pendingDecode, // 待解码包数
|
||||
pendingPlay, // 待播放包数
|
||||
totalPending: pendingDecode + pendingPlay // 总待处理包数
|
||||
};
|
||||
}
|
||||
|
||||
// 清空所有音频缓冲并停止播放
|
||||
clearAllAudio() {
|
||||
log('AudioPlayer: 清空所有音频', 'info');
|
||||
|
||||
// 清空接收队列(使用clear方法保持对象引用)
|
||||
this.queue.clear();
|
||||
|
||||
// 清空流上下文的所有缓冲
|
||||
if (this.streamingContext) {
|
||||
this.streamingContext.clearAllBuffers();
|
||||
}
|
||||
|
||||
log('AudioPlayer: 音频已清空', 'success');
|
||||
}
|
||||
}
|
||||
|
||||
// 创建单例
|
||||
|
||||
@@ -22,6 +22,7 @@ export class StreamingContext {
|
||||
this.source = null; // 当前音频源
|
||||
this.totalSamples = 0; // 累积的总样本数
|
||||
this.lastPlayTime = 0; // 上次播放的时间戳
|
||||
this.scheduledEndTime = 0; // 已调度音频的结束时间
|
||||
}
|
||||
|
||||
// 缓存音频数组
|
||||
@@ -31,17 +32,19 @@ export class StreamingContext {
|
||||
|
||||
// 获取需要处理缓存队列,单线程:在audioBufferQueue一直更新的状态下不会出现安全问题
|
||||
async getPendingAudioBufferQueue() {
|
||||
// 原子交换 + 清空
|
||||
[this.pendingAudioBufferQueue, this.audioBufferQueue] = [await this.audioBufferQueue.dequeue(), new BlockingQueue()];
|
||||
// 等待数据到达并获取
|
||||
const data = await this.audioBufferQueue.dequeue();
|
||||
// 赋值给待处理队列
|
||||
this.pendingAudioBufferQueue = data;
|
||||
}
|
||||
|
||||
// 获取正在播放已解码的PCM队列,单线程:在activeQueue一直更新的状态下不会出现安全问题
|
||||
async getQueue(minSamples) {
|
||||
let TepArray = [];
|
||||
const num = minSamples - this.queue.length > 0 ? minSamples - this.queue.length : 1;
|
||||
// 原子交换 + 清空
|
||||
[TepArray, this.activeQueue] = [await this.activeQueue.dequeue(num), new BlockingQueue()];
|
||||
this.queue.push(...TepArray);
|
||||
|
||||
// 等待数据并获取
|
||||
const tempArray = await this.activeQueue.dequeue(num);
|
||||
this.queue.push(...tempArray);
|
||||
}
|
||||
|
||||
// 将Int16音频数据转换为Float32音频数据
|
||||
@@ -54,6 +57,57 @@ export class StreamingContext {
|
||||
return float32Data;
|
||||
}
|
||||
|
||||
// 获取待解码包数
|
||||
getPendingDecodeCount() {
|
||||
return this.audioBufferQueue.length + this.pendingAudioBufferQueue.length;
|
||||
}
|
||||
|
||||
// 获取待播放样本数(转换为包数,每包960样本)
|
||||
getPendingPlayCount() {
|
||||
// 计算已在队列中的样本
|
||||
const queuedSamples = this.activeQueue.length + this.queue.length;
|
||||
|
||||
// 计算已调度但未播放的样本(在Web Audio缓冲区中)
|
||||
let scheduledSamples = 0;
|
||||
if (this.playing && this.scheduledEndTime) {
|
||||
const currentTime = this.audioContext.currentTime;
|
||||
const remainingTime = Math.max(0, this.scheduledEndTime - currentTime);
|
||||
scheduledSamples = Math.floor(remainingTime * this.sampleRate);
|
||||
}
|
||||
|
||||
const totalSamples = queuedSamples + scheduledSamples;
|
||||
return Math.ceil(totalSamples / 960);
|
||||
}
|
||||
|
||||
// 清空所有音频缓冲
|
||||
clearAllBuffers() {
|
||||
log('清空所有音频缓冲', 'info');
|
||||
|
||||
// 清空所有队列(使用clear方法保持对象引用)
|
||||
this.audioBufferQueue.clear();
|
||||
this.pendingAudioBufferQueue = [];
|
||||
this.activeQueue.clear();
|
||||
this.queue = [];
|
||||
|
||||
// 停止当前播放的音频源
|
||||
if (this.source) {
|
||||
try {
|
||||
this.source.stop();
|
||||
this.source.disconnect();
|
||||
} catch (e) {
|
||||
// 忽略已经停止的错误
|
||||
}
|
||||
this.source = null;
|
||||
}
|
||||
|
||||
// 重置状态
|
||||
this.playing = false;
|
||||
this.scheduledEndTime = this.audioContext.currentTime;
|
||||
this.totalSamples = 0;
|
||||
|
||||
log('音频缓冲已清空', 'success');
|
||||
}
|
||||
|
||||
// 将Opus数据解码为PCM
|
||||
async decodeOpusFrames() {
|
||||
if (!this.opusDecoder) {
|
||||
@@ -97,7 +151,7 @@ export class StreamingContext {
|
||||
|
||||
// 开始播放音频
|
||||
async startPlaying() {
|
||||
let scheduledEndTime = this.audioContext.currentTime; // 跟踪已调度音频的结束时间
|
||||
this.scheduledEndTime = this.audioContext.currentTime; // 跟踪已调度音频的结束时间
|
||||
|
||||
while (true) {
|
||||
// 初始缓冲:等待足够的样本再开始播放
|
||||
@@ -126,7 +180,7 @@ export class StreamingContext {
|
||||
|
||||
// 精确调度播放时间
|
||||
const currentTime = this.audioContext.currentTime;
|
||||
const startTime = Math.max(scheduledEndTime, currentTime);
|
||||
const startTime = Math.max(this.scheduledEndTime, currentTime);
|
||||
|
||||
// 直接连接到输出
|
||||
this.source.connect(this.audioContext.destination);
|
||||
@@ -136,7 +190,7 @@ export class StreamingContext {
|
||||
|
||||
// 更新下一个音频块的调度时间
|
||||
const duration = audioBuffer.duration;
|
||||
scheduledEndTime = startTime + duration;
|
||||
this.scheduledEndTime = startTime + duration;
|
||||
this.lastPlayTime = startTime;
|
||||
|
||||
// 如果队列中数据不足,等待新数据
|
||||
|
||||
@@ -123,7 +123,12 @@ export class WebSocketHandler {
|
||||
} else if (message.state === 'sentence_end') {
|
||||
log(`语音段结束: ${message.text}`, 'info');
|
||||
} else if (message.state === 'stop') {
|
||||
log('服务器语音传输结束', 'info');
|
||||
log('服务器语音传输结束,清空所有音频缓冲', 'info');
|
||||
|
||||
// 清空所有音频缓冲并停止播放
|
||||
const audioPlayer = getAudioPlayer();
|
||||
audioPlayer.clearAllAudio();
|
||||
|
||||
this.isRemoteSpeaking = false;
|
||||
if (this.onRecordButtonStateChange) {
|
||||
this.onRecordButtonStateChange(false);
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
// UI控制模块
|
||||
import { loadConfig, saveConfig } from '../config/manager.js';
|
||||
import { getAudioPlayer } from '../core/audio/player.js';
|
||||
import { getAudioRecorder } from '../core/audio/recorder.js';
|
||||
import { getWebSocketHandler } from '../core/network/websocket.js';
|
||||
|
||||
@@ -9,6 +10,7 @@ export class UIController {
|
||||
this.isEditing = false;
|
||||
this.visualizerCanvas = null;
|
||||
this.visualizerContext = null;
|
||||
this.audioStatsTimer = null;
|
||||
}
|
||||
|
||||
// 初始化
|
||||
@@ -18,6 +20,7 @@ export class UIController {
|
||||
|
||||
this.initVisualizer();
|
||||
this.initEventListeners();
|
||||
this.startAudioStatsMonitor();
|
||||
loadConfig();
|
||||
}
|
||||
|
||||
@@ -86,17 +89,20 @@ export class UIController {
|
||||
const sessionStatus = document.getElementById('sessionStatus');
|
||||
if (!sessionStatus) return;
|
||||
|
||||
// 保留背景元素
|
||||
const bgHtml = '<span id="sessionStatusBg" style="position: absolute; left: 0; top: 0; bottom: 0; width: 0%; background: linear-gradient(90deg, rgba(76, 175, 80, 0.2), rgba(33, 150, 243, 0.2)); transition: width 0.15s ease-out, background 0.3s ease; z-index: 0; border-radius: 20px;"></span>';
|
||||
|
||||
if (isSpeaking === null) {
|
||||
// 离线状态
|
||||
sessionStatus.innerHTML = '<span class="emoji-large">😶</span> 小智离线中';
|
||||
sessionStatus.innerHTML = bgHtml + '<span style="position: relative; z-index: 1;"><span class="emoji-large">😶</span> 小智离线中</span>';
|
||||
sessionStatus.className = 'status offline';
|
||||
} else if (isSpeaking) {
|
||||
// 说话中
|
||||
sessionStatus.innerHTML = '<span class="emoji-large">😶</span> 小智说话中';
|
||||
sessionStatus.innerHTML = bgHtml + '<span style="position: relative; z-index: 1;"><span class="emoji-large">😶</span> 小智说话中</span>';
|
||||
sessionStatus.className = 'status speaking';
|
||||
} else {
|
||||
// 聆听中
|
||||
sessionStatus.innerHTML = '<span class="emoji-large">😶</span> 小智聆听中';
|
||||
sessionStatus.innerHTML = bgHtml + '<span style="position: relative; z-index: 1;"><span class="emoji-large">😶</span> 小智聆听中</span>';
|
||||
sessionStatus.className = 'status listening';
|
||||
}
|
||||
}
|
||||
@@ -110,8 +116,72 @@ export class UIController {
|
||||
let currentText = sessionStatus.textContent;
|
||||
// 移除现有的表情符号
|
||||
currentText = currentText.replace(/[\u{1F300}-\u{1F9FF}]|[\u{2600}-\u{26FF}]|[\u{2700}-\u{27BF}]/gu, '').trim();
|
||||
|
||||
// 保留背景元素
|
||||
const bgHtml = '<span id="sessionStatusBg" style="position: absolute; left: 0; top: 0; bottom: 0; width: 0%; background: linear-gradient(90deg, rgba(76, 175, 80, 0.2), rgba(33, 150, 243, 0.2)); transition: width 0.15s ease-out, background 0.3s ease; z-index: 0; border-radius: 20px;"></span>';
|
||||
|
||||
// 使用 innerHTML 添加带样式的表情
|
||||
sessionStatus.innerHTML = `<span class="emoji-large">${emoji}</span> ${currentText}`;
|
||||
sessionStatus.innerHTML = bgHtml + `<span style="position: relative; z-index: 1;"><span class="emoji-large">${emoji}</span> ${currentText}</span>`;
|
||||
}
|
||||
|
||||
// 更新音频统计信息
|
||||
updateAudioStats() {
|
||||
const audioPlayer = getAudioPlayer();
|
||||
const stats = audioPlayer.getAudioStats();
|
||||
|
||||
const sessionStatus = document.getElementById('sessionStatus');
|
||||
const sessionStatusBg = document.getElementById('sessionStatusBg');
|
||||
|
||||
// 只在说话状态下显示背景进度
|
||||
if (sessionStatus && sessionStatus.classList.contains('speaking') && sessionStatusBg) {
|
||||
if (stats.pendingPlay > 0) {
|
||||
// 计算进度:5包=50%,10包及以上=100%
|
||||
let percentage;
|
||||
if (stats.pendingPlay >= 10) {
|
||||
percentage = 100;
|
||||
} else {
|
||||
percentage = (stats.pendingPlay / 10) * 100;
|
||||
}
|
||||
|
||||
sessionStatusBg.style.width = `${percentage}%`;
|
||||
|
||||
// 根据缓冲量改变背景颜色
|
||||
if (stats.pendingPlay < 5) {
|
||||
// 缓冲不足:橙红色半透明
|
||||
sessionStatusBg.style.background = 'linear-gradient(90deg, rgba(255, 152, 0, 0.25), rgba(255, 87, 34, 0.25))';
|
||||
} else if (stats.pendingPlay < 10) {
|
||||
// 一般:黄绿色半透明
|
||||
sessionStatusBg.style.background = 'linear-gradient(90deg, rgba(205, 220, 57, 0.25), rgba(76, 175, 80, 0.25))';
|
||||
} else {
|
||||
// 充足:绿蓝色半透明
|
||||
sessionStatusBg.style.background = 'linear-gradient(90deg, rgba(76, 175, 80, 0.25), rgba(33, 150, 243, 0.25))';
|
||||
}
|
||||
} else {
|
||||
// 没有缓冲,隐藏背景
|
||||
sessionStatusBg.style.width = '0%';
|
||||
}
|
||||
} else {
|
||||
// 非说话状态,隐藏背景
|
||||
if (sessionStatusBg) {
|
||||
sessionStatusBg.style.width = '0%';
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// 启动音频统计监控
|
||||
startAudioStatsMonitor() {
|
||||
// 每100ms更新一次音频统计
|
||||
this.audioStatsTimer = setInterval(() => {
|
||||
this.updateAudioStats();
|
||||
}, 100);
|
||||
}
|
||||
|
||||
// 停止音频统计监控
|
||||
stopAudioStatsMonitor() {
|
||||
if (this.audioStatsTimer) {
|
||||
clearInterval(this.audioStatsTimer);
|
||||
this.audioStatsTimer = null;
|
||||
}
|
||||
}
|
||||
|
||||
// 绘制音频可视化效果
|
||||
@@ -196,7 +266,6 @@ export class UIController {
|
||||
const deviceMacInput = document.getElementById('deviceMac');
|
||||
const deviceNameInput = document.getElementById('deviceName');
|
||||
const clientIdInput = document.getElementById('clientId');
|
||||
const tokenInput = document.getElementById('token');
|
||||
|
||||
toggleButton.addEventListener('click', () => {
|
||||
this.isEditing = !this.isEditing;
|
||||
@@ -204,7 +273,6 @@ export class UIController {
|
||||
deviceMacInput.disabled = !this.isEditing;
|
||||
deviceNameInput.disabled = !this.isEditing;
|
||||
clientIdInput.disabled = !this.isEditing;
|
||||
tokenInput.disabled = !this.isEditing;
|
||||
|
||||
toggleButton.textContent = this.isEditing ? '确定' : '编辑';
|
||||
|
||||
|
||||
@@ -95,4 +95,9 @@ export default class BlockingQueue {
|
||||
get length() {
|
||||
return this.#items.length;
|
||||
}
|
||||
|
||||
/* 清空队列(保持对象引用,不影响等待者) */
|
||||
clear() {
|
||||
this.#items.length = 0;
|
||||
}
|
||||
}
|
||||
@@ -52,20 +52,16 @@
|
||||
<label for="deviceMac">设备MAC:</label>
|
||||
<input type="text" id="deviceMac" placeholder="device-id" disabled>
|
||||
</div>
|
||||
<div class="config-item">
|
||||
<label for="deviceName">设备名称:</label>
|
||||
<input type="text" id="deviceName" value="Web测试设备" placeholder="deviceName" disabled>
|
||||
</div>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<div class="config-item">
|
||||
<label for="clientId">客户端ID:</label>
|
||||
<input type="text" id="clientId" value="web_test_client" placeholder="client-id"
|
||||
disabled>
|
||||
</div>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<div class="config-item">
|
||||
<label for="token">认证令牌:</label>
|
||||
<input type="text" id="token" value="your-token1" placeholder="token" disabled>
|
||||
<label for="deviceName">设备名称:</label>
|
||||
<input type="text" id="deviceName" value="Web测试设备" placeholder="deviceName" disabled>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
@@ -91,6 +87,7 @@
|
||||
<input type="text" id="serverUrl" value="" readonly disabled placeholder="点击连接按钮后,自动从OTA接口获取" />
|
||||
</div>
|
||||
</div>
|
||||
|
||||
</div>
|
||||
|
||||
<div class="section">
|
||||
@@ -113,9 +110,13 @@
|
||||
<canvas id="audioVisualizer" class="audio-visualizer"></canvas>
|
||||
</div>
|
||||
|
||||
<div style="margin: -10px 0px 5px 0px;">
|
||||
<div style="margin: -10px 0px 10px 0px;">
|
||||
<span class="connection-status llm-emoji">
|
||||
<span id="sessionStatus" class="status offline"><span class="emoji-large">😶</span> 小智离线中</span>
|
||||
<span id="sessionStatus" class="status offline" style="position: relative; display: inline-block;">
|
||||
<span id="sessionStatusBg"
|
||||
style="position: absolute; left: 0; top: 0; bottom: 0; width: 0%; background: linear-gradient(90deg, rgba(76, 175, 80, 0.2), rgba(33, 150, 243, 0.2)); transition: width 0.15s ease-out, background 0.3s ease; z-index: 0; border-radius: 20px;"></span>
|
||||
<span style="position: relative; z-index: 1;"><span class="emoji-large">😶</span> 小智离线中</span>
|
||||
</span>
|
||||
</span>
|
||||
</div>
|
||||
<div class="flex-container">
|
||||
|
||||