by waybarrios · MCP 服务器 · ★ 1.6k
vLLM-MLX vLLM-like inference for Apple Silicon - GPU-accelerated Text, Image, Video & Audio on Mac Overview vllm-mlx brings native Apple Silicon GPU acceleration to vLLM by integrating: MLX: Apple's ML framework with unified memory and Metal kernels mlx-lm: Optimized LLM inference with KV cache and quantization mlx-vlm: Vision-language models for multimodal inference mlx-audio: Speech-to-Text and Text-to-Speech with native voices mlx-embeddings: Text embeddings for semantic search and RAG Features Multimodal - Text, Image, Video & Audio in one platform Native GPU acceleration on Apple Silicon...
| Stars | 1,560 |
| Forks | 221 |
| Language | Python |
| Category | MCP 服务器 |
| License | Apache-2.0 |
| Quality Score | 70.9665687613498/100 |
| Open Issues | 98 |
| Last Updated | 2026-09-06 |
| Created | 2025-12-06 |
| Platforms | claude-code, mcp, python |
| Est. Tokens | ~16k |
These tools work well together with vllm-mlx for enhanced workflows:
Explore other popular mcp 服务器 tools:
vllm-mlx is High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.. It is categorized as a MCP 服务器 with 1.6k GitHub stars.
vllm-mlx is primarily written in Python. It covers topics such as anthropic, anthropic-api, apple-silicon.
You can find installation instructions and usage details in the vllm-mlx GitHub repository at github.com/waybarrios/vllm-mlx. The project has 1.6k stars and 221 forks, indicating an active community.
vllm-mlx is released under the Apache-2.0 license, making it free to use and modify according to the license terms.