doubao-asr Agent Skill for transcribing audio files via ByteDance Volcengine Seed-ASR 2.0 (豆包录音文件识别模型2.0) Best-in-class Chinese speech recognition — Mandarin, Cantonese, Sichuan dialect, and 13+ languages. With speaker diarization, up to 5 hours / 512MB per file. 中文语音识别准确率业界领先——支持普通话、粤语、四川话等方言及 13+ 种语言。自带说话人分离,单文件最长 5 小时 / 512MB。 Why this skill? Setting up Volcengine's Doubao Audio File Recognition 2.0 (豆包录音文件识别模型2.0, recorded audio → text) from scratch involves 4 environment variables across 3 different console pages (Speech console, IAM, TOS).
| Stars | 7 |
| Forks | 1 |
| Language | Shell |
| Category | AI 工具 |
| License | Apache-2.0 |
| Quality Score | 40.15/100 |
| Open Issues | 1 |
| Last Updated | 2026-06-10 |
| Created | 2026-02-28 |
| Platforms | claude-code, cli |
| Est. Tokens | ~3k |
Explore other popular ai 工具 tools:
doubao-asr is Agent Skill: Transcribe audio files via ByteDance Volcengine Seed-ASR 2.0 (豆包录音文件识别模型2.0). Best-in-class Chinese speech recognition with step-by-step setup guide.. It is categorized as a AI 工具 with 7 GitHub stars.
doubao-asr is primarily written in Shell. It covers topics such as agent-skill, asr, bytedance.
You can find installation instructions and usage details in the doubao-asr GitHub repository at github.com/vahnxu/doubao-asr. The project has 7 stars and 1 forks, indicating an active community.
doubao-asr is released under the Apache-2.0 license, making it free to use and modify according to the license terms.