Files
aiagent/docs/产品化落地方案.md
renjianbo a7512a5423 docs: update 3 reference docs + add Android design, Feishu bot config, productization plan
- Rewrite api-reference.md: 245 endpoints across 38 modules, correct auth paths and response format
- Rewrite 内置工具列表.md: all 56 real tools in 11 categories
- Fix quickstart.md: local dev ports (3001/8038) vs Docker (8037/8038), --port 8038
- Add android-app-design.md: Kotlin/Compose/MVVM design with SSE, FCM, voice
- Add 飞书智能体配置手册.md: all 6 bots config, capabilities, memory architecture
- Add 产品化落地方案.md: PWA/voice/push/Flutter productization roadmap

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-13 21:58:10 +08:00

15 KiB
Raw Blame History

豆包风格智能助手 — 产品化落地方案

从 AI Agent 到完整产品:App、语音、推送三大能力落地路线图


一、现状盘点

能力 后端工具 前端 状态
语音转文字 speech_to_text (API:OpenAI Whisper) 无录音入口 ⚠️ 工具存在,前端缺失
文字转语音 text_to_speech (API:OpenAI TTS) 无播放器 ⚠️ 工具存在,前端缺失
通知推送 notify_user (Main Agent) 无推送UI ⚠️ 工具存在,前端缺失
消息通知 Notification 模型 + API 无 ⚠️ DB就绪,无消费端
移动App — 无 ❌ 不存在
浏览器推送 — 无 ❌ Service Worker 未配置
飞书消息 飞书 Bot 长连接 — ✅ 已有

二、App 客户端方案

2.1 技术选型

方案 优势 劣势 推荐度
Flutter 一套代码双端、热重载、Dart易学 包体积较大 ⭐⭐⭐⭐⭐
React Native JS生态、热更新、社区大 原生交互需桥接 ⭐⭐⭐⭐
PWA (渐进式Web) 零安装、复用Vue代码、成本最低 功能受限(推送iOS) ⭐⭐⭐⭐
UniApp 国内小程序+App一体 性能不如Flutter ⭐⭐⭐

推荐路径:PWA 先行(1周上线),Flutter 跟进(4周MVP)

2.2 PWA 快速上线方案(第1周)

frontend/
├── public/
│   ├── manifest.json          # PWA 配置(图标/名称/主题色)
│   └── sw.js                  # Service Worker(离线缓存+推送)
├── src/
│   ├── views/
│   │   └── MobileChat.vue     # 移动端对话界面
│   └── utils/
│       └── push.ts            # 浏览器推送注册

manifest.json 示例:

{
  "name": "豆包智能助手",
  "short_name": "豆包",
  "start_url": "/?source=pwa",
  "display": "standalone",
  "background_color": "#ffffff",
  "theme_color": "#4f46e5",
  "icons": [
    {"src": "/icons/icon-192.png", "sizes": "192x192", "type": "image/png"},
    {"src": "/icons/icon-512.png", "sizes": "512x512", "type": "image/png"}
  ]
}

核心改动:

  1. vite.config.ts 增加 @vitejs/plugin-pwa 插件
  2. 新建 MobileChat.vue — 移动端全屏聊天界面(底部输入框+语音按钮+对话气泡)
  3. 新建 src/utils/push.ts — 注册 Service Worker、请求通知权限
  4. index.html 添加 <link rel="manifest"> 和 <meta name="theme-color">

代价: ~200行新代码,前端依赖 +1 (vite-plugin-pwa)

2.3 Flutter App 方案(第2-5周)

doubao_app/
├── lib/
│   ├── main.dart                   # 入口
│   ├── app.dart                    # MaterialApp 配置
│   ├── models/
│   │   ├── message.dart            # 消息模型
│   │   └── user.dart               # 用户模型
│   ├── services/
│   │   ├── api_service.dart        # HTTP 客户端(复用后端API)
│   │   ├── auth_service.dart       # JWT 存储/刷新
│   │   ├── audio_service.dart      # 录音 + 播放
│   │   └── push_service.dart       # FCM/个推 注册
│   ├── pages/
│   │   ├── login_page.dart         # 登录
│   │   ├── chat_page.dart          # 对话主界面
│   │   ├── history_page.dart       # 历史对话
│   │   └── settings_page.dart      # 设置
│   └── widgets/
│       ├── chat_bubble.dart        # 对话气泡
│       ├── voice_button.dart       # 语音录制按钮
│       └── typing_indicator.dart   # 输入状态
├── pubspec.yaml
└── README.md

核心功能实现:

// lib/services/api_service.dart
class ApiService {
  static const baseUrl = 'http://101.43.95.130:8038/api/v1';
  
  // 流式对话(SSE)
  Stream<String> chatStream(String agentId, String message) async* {
    final response = await http.Client().send(
      http.Request('POST', Uri.parse('$baseUrl/agent-chat/$agentId/stream'))
        ..headers.addAll({'Authorization': 'Bearer $token', 'Content-Type': 'application/json'})
        ..body = jsonEncode({'message': message, 'streamlined': true}),
    );
    await for (final chunk in response.stream.transform(utf8.decoder)) {
      // 解析 SSE data: {...} 事件
      yield chunk;
    }
  }
}

代价: ~1500行 Dart 代码,Flutter SDK + 依赖(dio, flutter_secure_storage, record, audioplayers, firebase_messaging)


三、语音能力方案

3.1 需求拆解

语音输入(ASR)                             语音输出(TTS)
┌──────────────────────┐                   ┌──────────────────┐
│ 用户说话              │                   │ AI回复文本        │
│   ↓                   │                   │   ↓              │
│ 前端录音 → Base64     │                   │ 后端TTS → .mp3   │
│   ↓                   │                   │   ↓              │
│ 后端 speech_to_text   │                   │ 返回音频URL       │
│   ↓                   │                   │   ↓              │
│ Whisper API 转文字    │                   │ 前端播放器        │
│   ↓                   │                   │                  │
│ 送入Agent对话         │                   │                  │
└──────────────────────┘                   └──────────────────┘

3.2 前端录音实现(Vue3)

// src/composables/useVoiceInput.ts
export function useVoiceInput() {
  const isRecording = ref(false)
  let mediaRecorder: MediaRecorder | null = null
  const chunks: Blob[] = []

  async function startRecording() {
    const stream = await navigator.mediaDevices.getUserMedia({ audio: true })
    mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/webm' })
    mediaRecorder.ondataavailable = (e) => chunks.push(e.data)
    mediaRecorder.start()
    isRecording.value = true
  }

  async function stopRecording(): Promise<Blob> {
    return new Promise((resolve) => {
      mediaRecorder!.onstop = () => {
        resolve(new Blob(chunks, { type: 'audio/webm' }))
        chunks.length = 0
      }
      mediaRecorder!.stop()
      mediaRecorder!.stream.getTracks().forEach(t => t.stop())
      isRecording.value = false
    })
  }

  return { isRecording, startRecording, stopRecording }
}

3.3 后端语音端点(新增)

# app/api/voice.py — 新增文件
@router.post("/voice/asr")
async def voice_to_text(
    file: UploadFile = File(...),
    current_user: User = Depends(get_current_user),
):
    """语音转文字 — 前端录音上传 → Whisper → 返回文本"""
    audio_path = f"/tmp/{uuid4()}.webm"
    with open(audio_path, "wb") as f:
        f.write(await file.read())
    text = await speech_to_text_tool(audio_path)
    os.remove(audio_path)
    return {"text": text}

@router.post("/voice/tts")
async def text_to_voice(
    req: TTSRequest,
    current_user: User = Depends(get_current_user),
):
    """文字转语音 — 返回音频文件URL"""
    output_path = f"uploads/tts/{uuid4()}.mp3"
    result = await text_to_speech_tool(req.text, req.voice, output_path)
    return {"audio_url": f"/api/v1/uploads/tts/{output_path}"}

代价: 后端 ~80行新增,前端 ~60行 composable

3.4 语音交互优化

优化点 方案
流式TTS OpenAI TTS 不支持流式,可换用 Edge-TTS(免费)或 ElevenLabs 流式API
VAD静音检测 @ricky0123/vad-web — 前端自动检测说话结束,无需手动停止
打断对话 用户开始说话时中止当前TTS播放 + 打断LLM流式输出
音色选择 提供 6 种音色(alloy/echo/fable/onyx/nova/shimmer)切换

四、推送通知方案

4.1 架构

┌─────────────────┐     ┌──────────────┐     ┌──────────────────┐
│ Agent 完成任务    │ →   │ Notification │  →  │ FCM / 个推        │
│ schedule 触发     │     │ DB 表 + API  │     │ Push Service     │
│ notify_user 工具  │     │              │     │         ↓         │
└─────────────────┘     └──────────────┘     │  手机/浏览器       │
                                             └──────────────────┘

4.2 通知类型定义

类型 触发场景 优先级
agent_reply Agent 完成回复 中
schedule_done 定时任务执行完毕 中
goal_milestone 目标达成阶段性成果 高
approval_required 工具调用需要人类审批 紧急
alert 系统告警触发 紧急
daily_summary 每日 AI 摘要推送 低

4.3 浏览器推送(PWA)

// public/sw.js
self.addEventListener('push', (event) => {
  const data = event.data?.json() || {}
  self.registration.showNotification(data.title, {
    body: data.body,
    icon: '/icons/icon-192.png',
    badge: '/icons/badge-72.png',
    data: { url: data.url || '/' },
    actions: data.actions || [],
    requireInteraction: data.priority === 'urgent',
  })
})

self.addEventListener('notificationclick', (event) => {
  event.notification.close()
  event.waitUntil(clients.openWindow(event.notification.data.url))
})
// src/utils/push.ts
export async function subscribeToPush(): Promise<string | null> {
  const reg = await navigator.serviceWorker.ready
  const sub = await reg.pushManager.subscribe({
    userVisibleOnly: true,
    applicationServerKey: urlBase64ToUint8Array(VAPID_PUBLIC_KEY),
  })
  // 将 subscription 发送到后端 /api/v1/push/subscribe
  await api.post('/push/subscribe', { subscription: sub.toJSON() })
  return sub.endpoint
}

4.4 后端推送服务(新增)

# app/services/push_service.py — 新增文件
import json
from pywebpush import webpush, WebPushException

VAPID_CLAIMS = {
    "sub": "mailto:admin@tiangong.ai"
}

async def send_web_push(user_id: str, title: str, body: str, url: str = "/"):
    """向用户的所有浏览器端点推送通知"""
    subscriptions = await get_user_push_subscriptions(user_id)
    for sub in subscriptions:
        try:
            webpush(
                subscription_info=json.loads(sub.endpoint_data),
                data=json.dumps({"title": title, "body": body, "url": url}),
                vapid_private_key=VAPID_PRIVATE_KEY,
                vapid_claims=VAPID_CLAIMS,
            )
        except WebPushException:
            # 端点失效,标记删除
            await mark_subscription_expired(sub.id)

代价: 后端 ~120行,新增依赖 pywebpush,新增 push_subscriptions 表

4.5 App 推送(FCM/个推)

Flutter 侧接入 firebase_messaging:

// 注册 FCM token
final token = await FirebaseMessaging.instance.getToken();
await api.post('/push/register-app', {'token': token, 'platform': 'android'});

// 前台消息处理
FirebaseMessaging.onMessage.listen((RemoteMessage message) {
  if (message.notification != null) {
    showLocalNotification(message.notification!);
  }
});

Android 渠道配置:

渠道ID 名称 重要性 行为
agent_reply AI回复 DEFAULT 声音+状态栏
approval 需要审批 HIGH 声音+振动+悬浮
alert 系统告警 URGENT 全屏通知

五、落地路线图

Week 1                Week 2-3              Week 4-5              长期
┌─────────────────┐  ┌─────────────────┐  ┌─────────────────┐  ┌──────────────┐
│ ██ PWA 快速上线  │  │ ██ 语音交互      │  │ ██ Flutter App  │  │ 多模态输入    │
│                 │  │                 │  │                 │  │             │
│ • manifest.json │  │ • ASR 前端录音   │  │ • Flutter工程   │  │ • 图片上传    │
│ • ServiceWorker │  │ • TTS 播放器     │  │ • 对话主界面    │  │ • 文档分析    │
│ • MobileChat.vue│  │ • 语音API端点    │  │ • 推送接入FCM   │  │ • 拍照识别    │
│ • 浏览器推送    │  │ • VAD静音检测    │  │ • Android/iOS   │  │ • 视频理解    │
└─────────────────┘  └─────────────────┘  └─────────────────┘  └──────────────┘
  ~200行前端改动       ~200行前后端          ~1500行Flutter        ~500行
  1天                 3天                  2周                   按需

里程碑验收标准

里程碑 验收标准
M1: PWA上线 手机浏览器打开 → 添加到主屏幕 → 能对话 → 能收推送
M2: 语音可用 点击麦克风 → 说话 → 自动转文字 → AI回复可朗读
M3: App上架 Flutter App 在 TestFlight/应用宝 可下载安装
M4: 完整产品 拍照→AI分析 + 语音对话 + 推送通知 + 离线缓存

六、工作量估算

模块 内容 工作量
PWA manifest + SW + MobileChat + 推送注册 1天
语音ASR 前端录音composable + 后端voice/asr端点 0.5天
语音TTS 前端播放器 + 后端voice/tts端点 0.5天
浏览器推送 sw.js推送处理 + push.ts注册 + 后端推送服务 + push_subscriptions表 1天
Flutter App 登录/对话/历史/设置/推送/语音/图标/上架 2周
后端推送增强 VAPID配置 + FCM集成 + 通知策略引擎 1天
测试联调 端到端测试 + 兼容性测试 2天
合计 约3周

七、当前可用能力速查

以下工具后端已存在,前端接入即可用:

工具名 功能 前端需要 难度
speech_to_text Whisper 语音转文字 录音按钮 + 上传 ★
text_to_speech OpenAI TTS 文字转语音 播放按钮 + 音频播放 ★
notify_user Main Agent 通知 通知中心UI ★★
image_vision 图片视觉理解 图片上传组件 ★
image_ocr 图片文字提取 图片上传+结果展示 ★
deploy_push 部署推送 一键部署按钮 ★
feishu_* 飞书系列(7个工具) — ✅ 已可用

结论: 从 AI Agent 到完整产品,核心 AI 能力已经超越真豆包的部分维度。剩下的工作主要是 前端交互层 和 分发渠道,成本可控,3周可达 MVP 水平。