by tristan-mcinnis · AI 工具 · ★ 10
Multi-Modal AI Voice Assistant A multi-modal AI voice assistant supporting DeepSeek (default), OpenAI, Anthropic Claude, and local LM Studio LLMs with configurable text-to-speech (OpenAI streaming or Kokoro). Combines voice transcription, tool calling, clipboard extraction, screenshot analysis, and web search to respond with rich context.
| Stars | 10 |
| Forks | 3 |
| Language | Python |
| Category | AI 工具 |
| License | MIT |
| Quality Score | 40.3/100 |
| Last Updated | 2026-05-09 |
| Created | 2024-06-22 |
| Platforms | claude-code, cli, python |
| Est. Tokens | ~8k |
These tools work well together with Multimodal-voice-assistant for enhanced workflows:
Explore other popular ai 工具 tools:
Multimodal-voice-assistant is This project is a multi-modal AI voice assistant that uses LM Studio, OpenAI API or Claude Code, audio processing with WhisperModel, speech recognition, clipboard extraction, and image processing to r. It is categorized as a AI 工具 with 10 GitHub stars.
Multimodal-voice-assistant is primarily written in Python. It covers topics such as ai, ai-assistant, assistant.
You can find installation instructions and usage details in the Multimodal-voice-assistant GitHub repository at github.com/tristan-mcinnis/Multimodal-voice-assistant. The project has 10 stars and 3 forks, indicating an active community.
Multimodal-voice-assistant is released under the MIT license, making it free to use and modify according to the license terms.