- Rewrite api-reference.md: 245 endpoints across 38 modules, correct auth paths and response format - Rewrite 内置工具列表.md: all 56 real tools in 11 categories - Fix quickstart.md: local dev ports (3001/8038) vs Docker (8037/8038), --port 8038 - Add android-app-design.md: Kotlin/Compose/MVVM design with SSE, FCM, voice - Add 飞书智能体配置手册.md: all 6 bots config, capabilities, memory architecture - Add 产品化落地方案.md: PWA/voice/push/Flutter productization roadmap Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
15 KiB
15 KiB
豆包风格智能助手 — 产品化落地方案
从 AI Agent 到完整产品:App、语音、推送三大能力落地路线图
一、现状盘点
| 能力 | 后端工具 | 前端 | 状态 |
|---|---|---|---|
| 语音转文字 | speech_to_text (API:OpenAI Whisper) |
无录音入口 | ⚠️ 工具存在,前端缺失 |
| 文字转语音 | text_to_speech (API:OpenAI TTS) |
无播放器 | ⚠️ 工具存在,前端缺失 |
| 通知推送 | notify_user (Main Agent) |
无推送UI | ⚠️ 工具存在,前端缺失 |
| 消息通知 | Notification 模型 + API |
无 | ⚠️ DB就绪,无消费端 |
| 移动App | — | 无 | ❌ 不存在 |
| 浏览器推送 | — | 无 | ❌ Service Worker 未配置 |
| 飞书消息 | 飞书 Bot 长连接 | — | ✅ 已有 |
二、App 客户端方案
2.1 技术选型
| 方案 | 优势 | 劣势 | 推荐度 |
|---|---|---|---|
| Flutter | 一套代码双端、热重载、Dart易学 | 包体积较大 | ⭐⭐⭐⭐⭐ |
| React Native | JS生态、热更新、社区大 | 原生交互需桥接 | ⭐⭐⭐⭐ |
| PWA (渐进式Web) | 零安装、复用Vue代码、成本最低 | 功能受限(推送iOS) | ⭐⭐⭐⭐ |
| UniApp | 国内小程序+App一体 | 性能不如Flutter | ⭐⭐⭐ |
推荐路径:PWA 先行(1周上线),Flutter 跟进(4周MVP)
2.2 PWA 快速上线方案(第1周)
frontend/
├── public/
│ ├── manifest.json # PWA 配置(图标/名称/主题色)
│ └── sw.js # Service Worker(离线缓存+推送)
├── src/
│ ├── views/
│ │ └── MobileChat.vue # 移动端对话界面
│ └── utils/
│ └── push.ts # 浏览器推送注册
manifest.json 示例:
{
"name": "豆包智能助手",
"short_name": "豆包",
"start_url": "/?source=pwa",
"display": "standalone",
"background_color": "#ffffff",
"theme_color": "#4f46e5",
"icons": [
{"src": "/icons/icon-192.png", "sizes": "192x192", "type": "image/png"},
{"src": "/icons/icon-512.png", "sizes": "512x512", "type": "image/png"}
]
}
核心改动:
vite.config.ts增加@vitejs/plugin-pwa插件- 新建
MobileChat.vue— 移动端全屏聊天界面(底部输入框+语音按钮+对话气泡) - 新建
src/utils/push.ts— 注册 Service Worker、请求通知权限 index.html添加<link rel="manifest">和<meta name="theme-color">
代价: ~200行新代码,前端依赖 +1 (vite-plugin-pwa)
2.3 Flutter App 方案(第2-5周)
doubao_app/
├── lib/
│ ├── main.dart # 入口
│ ├── app.dart # MaterialApp 配置
│ ├── models/
│ │ ├── message.dart # 消息模型
│ │ └── user.dart # 用户模型
│ ├── services/
│ │ ├── api_service.dart # HTTP 客户端(复用后端API)
│ │ ├── auth_service.dart # JWT 存储/刷新
│ │ ├── audio_service.dart # 录音 + 播放
│ │ └── push_service.dart # FCM/个推 注册
│ ├── pages/
│ │ ├── login_page.dart # 登录
│ │ ├── chat_page.dart # 对话主界面
│ │ ├── history_page.dart # 历史对话
│ │ └── settings_page.dart # 设置
│ └── widgets/
│ ├── chat_bubble.dart # 对话气泡
│ ├── voice_button.dart # 语音录制按钮
│ └── typing_indicator.dart # 输入状态
├── pubspec.yaml
└── README.md
核心功能实现:
// lib/services/api_service.dart
class ApiService {
static const baseUrl = 'http://101.43.95.130:8038/api/v1';
// 流式对话(SSE)
Stream<String> chatStream(String agentId, String message) async* {
final response = await http.Client().send(
http.Request('POST', Uri.parse('$baseUrl/agent-chat/$agentId/stream'))
..headers.addAll({'Authorization': 'Bearer $token', 'Content-Type': 'application/json'})
..body = jsonEncode({'message': message, 'streamlined': true}),
);
await for (final chunk in response.stream.transform(utf8.decoder)) {
// 解析 SSE data: {...} 事件
yield chunk;
}
}
}
代价: ~1500行 Dart 代码,Flutter SDK + 依赖(dio, flutter_secure_storage, record, audioplayers, firebase_messaging)
三、语音能力方案
3.1 需求拆解
语音输入(ASR) 语音输出(TTS)
┌──────────────────────┐ ┌──────────────────┐
│ 用户说话 │ │ AI回复文本 │
│ ↓ │ │ ↓ │
│ 前端录音 → Base64 │ │ 后端TTS → .mp3 │
│ ↓ │ │ ↓ │
│ 后端 speech_to_text │ │ 返回音频URL │
│ ↓ │ │ ↓ │
│ Whisper API 转文字 │ │ 前端播放器 │
│ ↓ │ │ │
│ 送入Agent对话 │ │ │
└──────────────────────┘ └──────────────────┘
3.2 前端录音实现(Vue3)
// src/composables/useVoiceInput.ts
export function useVoiceInput() {
const isRecording = ref(false)
let mediaRecorder: MediaRecorder | null = null
const chunks: Blob[] = []
async function startRecording() {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true })
mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/webm' })
mediaRecorder.ondataavailable = (e) => chunks.push(e.data)
mediaRecorder.start()
isRecording.value = true
}
async function stopRecording(): Promise<Blob> {
return new Promise((resolve) => {
mediaRecorder!.onstop = () => {
resolve(new Blob(chunks, { type: 'audio/webm' }))
chunks.length = 0
}
mediaRecorder!.stop()
mediaRecorder!.stream.getTracks().forEach(t => t.stop())
isRecording.value = false
})
}
return { isRecording, startRecording, stopRecording }
}
3.3 后端语音端点(新增)
# app/api/voice.py — 新增文件
@router.post("/voice/asr")
async def voice_to_text(
file: UploadFile = File(...),
current_user: User = Depends(get_current_user),
):
"""语音转文字 — 前端录音上传 → Whisper → 返回文本"""
audio_path = f"/tmp/{uuid4()}.webm"
with open(audio_path, "wb") as f:
f.write(await file.read())
text = await speech_to_text_tool(audio_path)
os.remove(audio_path)
return {"text": text}
@router.post("/voice/tts")
async def text_to_voice(
req: TTSRequest,
current_user: User = Depends(get_current_user),
):
"""文字转语音 — 返回音频文件URL"""
output_path = f"uploads/tts/{uuid4()}.mp3"
result = await text_to_speech_tool(req.text, req.voice, output_path)
return {"audio_url": f"/api/v1/uploads/tts/{output_path}"}
代价: 后端 ~80行新增,前端 ~60行 composable
3.4 语音交互优化
| 优化点 | 方案 |
|---|---|
| 流式TTS | OpenAI TTS 不支持流式,可换用 Edge-TTS(免费)或 ElevenLabs 流式API |
| VAD静音检测 | @ricky0123/vad-web — 前端自动检测说话结束,无需手动停止 |
| 打断对话 | 用户开始说话时中止当前TTS播放 + 打断LLM流式输出 |
| 音色选择 | 提供 6 种音色(alloy/echo/fable/onyx/nova/shimmer)切换 |
四、推送通知方案
4.1 架构
┌─────────────────┐ ┌──────────────┐ ┌──────────────────┐
│ Agent 完成任务 │ → │ Notification │ → │ FCM / 个推 │
│ schedule 触发 │ │ DB 表 + API │ │ Push Service │
│ notify_user 工具 │ │ │ │ ↓ │
└─────────────────┘ └──────────────┘ │ 手机/浏览器 │
└──────────────────┘
4.2 通知类型定义
| 类型 | 触发场景 | 优先级 |
|---|---|---|
agent_reply |
Agent 完成回复 | 中 |
schedule_done |
定时任务执行完毕 | 中 |
goal_milestone |
目标达成阶段性成果 | 高 |
approval_required |
工具调用需要人类审批 | 紧急 |
alert |
系统告警触发 | 紧急 |
daily_summary |
每日 AI 摘要推送 | 低 |
4.3 浏览器推送(PWA)
// public/sw.js
self.addEventListener('push', (event) => {
const data = event.data?.json() || {}
self.registration.showNotification(data.title, {
body: data.body,
icon: '/icons/icon-192.png',
badge: '/icons/badge-72.png',
data: { url: data.url || '/' },
actions: data.actions || [],
requireInteraction: data.priority === 'urgent',
})
})
self.addEventListener('notificationclick', (event) => {
event.notification.close()
event.waitUntil(clients.openWindow(event.notification.data.url))
})
// src/utils/push.ts
export async function subscribeToPush(): Promise<string | null> {
const reg = await navigator.serviceWorker.ready
const sub = await reg.pushManager.subscribe({
userVisibleOnly: true,
applicationServerKey: urlBase64ToUint8Array(VAPID_PUBLIC_KEY),
})
// 将 subscription 发送到后端 /api/v1/push/subscribe
await api.post('/push/subscribe', { subscription: sub.toJSON() })
return sub.endpoint
}
4.4 后端推送服务(新增)
# app/services/push_service.py — 新增文件
import json
from pywebpush import webpush, WebPushException
VAPID_CLAIMS = {
"sub": "mailto:admin@tiangong.ai"
}
async def send_web_push(user_id: str, title: str, body: str, url: str = "/"):
"""向用户的所有浏览器端点推送通知"""
subscriptions = await get_user_push_subscriptions(user_id)
for sub in subscriptions:
try:
webpush(
subscription_info=json.loads(sub.endpoint_data),
data=json.dumps({"title": title, "body": body, "url": url}),
vapid_private_key=VAPID_PRIVATE_KEY,
vapid_claims=VAPID_CLAIMS,
)
except WebPushException:
# 端点失效,标记删除
await mark_subscription_expired(sub.id)
代价: 后端 ~120行,新增依赖 pywebpush,新增 push_subscriptions 表
4.5 App 推送(FCM/个推)
Flutter 侧接入 firebase_messaging:
// 注册 FCM token
final token = await FirebaseMessaging.instance.getToken();
await api.post('/push/register-app', {'token': token, 'platform': 'android'});
// 前台消息处理
FirebaseMessaging.onMessage.listen((RemoteMessage message) {
if (message.notification != null) {
showLocalNotification(message.notification!);
}
});
Android 渠道配置:
| 渠道ID | 名称 | 重要性 | 行为 |
|---|---|---|---|
agent_reply |
AI回复 | DEFAULT | 声音+状态栏 |
approval |
需要审批 | HIGH | 声音+振动+悬浮 |
alert |
系统告警 | URGENT | 全屏通知 |
五、落地路线图
Week 1 Week 2-3 Week 4-5 长期
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌──────────────┐
│ ██ PWA 快速上线 │ │ ██ 语音交互 │ │ ██ Flutter App │ │ 多模态输入 │
│ │ │ │ │ │ │ │
│ • manifest.json │ │ • ASR 前端录音 │ │ • Flutter工程 │ │ • 图片上传 │
│ • ServiceWorker │ │ • TTS 播放器 │ │ • 对话主界面 │ │ • 文档分析 │
│ • MobileChat.vue│ │ • 语音API端点 │ │ • 推送接入FCM │ │ • 拍照识别 │
│ • 浏览器推送 │ │ • VAD静音检测 │ │ • Android/iOS │ │ • 视频理解 │
└─────────────────┘ └─────────────────┘ └─────────────────┘ └──────────────┘
~200行前端改动 ~200行前后端 ~1500行Flutter ~500行
1天 3天 2周 按需
里程碑验收标准
| 里程碑 | 验收标准 |
|---|---|
| M1: PWA上线 | 手机浏览器打开 → 添加到主屏幕 → 能对话 → 能收推送 |
| M2: 语音可用 | 点击麦克风 → 说话 → 自动转文字 → AI回复可朗读 |
| M3: App上架 | Flutter App 在 TestFlight/应用宝 可下载安装 |
| M4: 完整产品 | 拍照→AI分析 + 语音对话 + 推送通知 + 离线缓存 |
六、工作量估算
| 模块 | 内容 | 工作量 |
|---|---|---|
| PWA | manifest + SW + MobileChat + 推送注册 | 1天 |
| 语音ASR | 前端录音composable + 后端voice/asr端点 | 0.5天 |
| 语音TTS | 前端播放器 + 后端voice/tts端点 | 0.5天 |
| 浏览器推送 | sw.js推送处理 + push.ts注册 + 后端推送服务 + push_subscriptions表 | 1天 |
| Flutter App | 登录/对话/历史/设置/推送/语音/图标/上架 | 2周 |
| 后端推送增强 | VAPID配置 + FCM集成 + 通知策略引擎 | 1天 |
| 测试联调 | 端到端测试 + 兼容性测试 | 2天 |
| 合计 | 约3周 |
七、当前可用能力速查
以下工具后端已存在,前端接入即可用:
| 工具名 | 功能 | 前端需要 | 难度 |
|---|---|---|---|
speech_to_text |
Whisper 语音转文字 | 录音按钮 + 上传 | ★ |
text_to_speech |
OpenAI TTS 文字转语音 | 播放按钮 + 音频播放 | ★ |
notify_user |
Main Agent 通知 | 通知中心UI | ★★ |
image_vision |
图片视觉理解 | 图片上传组件 | ★ |
image_ocr |
图片文字提取 | 图片上传+结果展示 | ★ |
deploy_push |
部署推送 | 一键部署按钮 | ★ |
feishu_* |
飞书系列(7个工具) | — | ✅ 已可用 |
结论: 从 AI Agent 到完整产品,核心 AI 能力已经超越真豆包的部分维度。剩下的工作主要是 前端交互层 和 分发渠道,成本可控,3周可达 MVP 水平。