Multimodal-voice-assistant

by tristan-mcinnis · AI 工具 · ★ 10

About Multimodal-voice-assistant

Multi-Modal AI Voice Assistant A multi-modal AI voice assistant supporting DeepSeek (default), OpenAI, Anthropic Claude, and local LM Studio LLMs with configurable text-to-speech (OpenAI streaming or Kokoro). Combines voice transcription, tool calling, clipboard extraction, screenshot analysis, and web search to respond with rich context.

aiai-assistantassistantclaude-codeimage-processingllmlm-studiomacosmultimodalopenai

Quick Facts

Stars10
Forks3
LanguagePython
CategoryAI 工具
LicenseMIT
Quality Score40.3/100
Last Updated2026-05-09
Created2024-06-22
Platformsclaude-code, cli, python
Est. Tokens~8k

Compatible Skills

These tools work well together with Multimodal-voice-assistant for enhanced workflows:

  • Whisper-Skill — semantic(0.17)+complementary+shared_fw(openai)+rare_topics+same_lang+similar_pop+shared_platform (68%)
  • vllm-mlx — semantic(0.24)+complementary+shared_fw(openai)+rare_topics+same_lang+shared_platform (65%)
  • pixelpanda-mcp — semantic(0.31)+complementary+rare_topics+same_lang+similar_pop+shared_platform (65%)
  • houtini-lm — semantic(0.35)+complementary+shared_fw(openai)+rare_topics+similar_pop+shared_platform (65%)
  • claude-code-tts — semantic(0.31)+complementary+shared_fw(openai)+rare_topics+similar_pop+shared_platform (63%)

More AI 工具 Tools

Explore other popular ai 工具 tools:

View all AI 工具 tools →

Popular Python Agent Tools

Frequently Asked Questions

What is Multimodal-voice-assistant?

Multimodal-voice-assistant is This project is a multi-modal AI voice assistant that uses LM Studio, OpenAI API or Claude Code, audio processing with WhisperModel, speech recognition, clipboard extraction, and image processing to r. It is categorized as a AI 工具 with 10 GitHub stars.

What programming language is Multimodal-voice-assistant written in?

Multimodal-voice-assistant is primarily written in Python. It covers topics such as ai, ai-assistant, assistant.

How do I install or use Multimodal-voice-assistant?

You can find installation instructions and usage details in the Multimodal-voice-assistant GitHub repository at github.com/tristan-mcinnis/Multimodal-voice-assistant. The project has 10 stars and 3 forks, indicating an active community.

What license does Multimodal-voice-assistant use?

Multimodal-voice-assistant is released under the MIT license, making it free to use and modify according to the license terms.

View on GitHub → Browse AI 工具 tools