Repo of the Day
hwdsl2/docker-ai-stack: Deploy a complete, self-hosted AI stack on your own server with one command. Includes Ollama (LLM), LiteLLM (AI gateway), Whisper (STT), Kokoro (TTS), Embeddings (RAG), and MCP Gateway. Most services run locally; LiteLLM optionall…
Published: Oct 10, 2026
Open repository ↗Deploy a complete, self-hosted AI stack on your own server with one command. Includes Ollama (LLM), LiteLLM (AI gateway), Whisper (STT), Kokoro (TTS), Embeddings (RAG), and MCP Gateway. Most servic...
Summary
hwdsl2/self-hosted-ai-stack is a Docker Compose project that bundles local LLM serving, chat, document parsing, embeddings, speech-to-text, text-to-speech, and MCP tool connectivity into one deployable system. It wires Ollama, LiteLLM, AnythingLLM, Hugging Face TEI, Whisper, Kokoro, Docling, and an MCP hub together with auto-generated credentials. The project targets engineers who want a self-hosted AI platform running locally on Linux (amd64 or arm64) with optional NVIDIA CUDA acceleration.
What it is useful for
The stack is useful when you want to run AI workloads against your own infrastructure instead of third-party APIs:
- Local LLM inference via InferCrate (Ollama-based, port 11434), supporting models like
llama3.2:3b, with requests routed through GatewayCrate (LiteLLM) on port 4000. - Web chat UI through AnythingLLM (port 3001), password-protected by default with an auto-generated admin password stored in the
anythingllm-datavolume. - Embeddings and RAG via EmbedCrate (port 8000) plus a pgvector-enabled Postgres, so semantic search works without a separate vector database.
- Voice pipelines using ScribeCrate for transcription, SpeakCrate for TTS, and ScribeCrate Live for real-time WebSocket STT over WebSocket.
- MCP tool access via ToolUplink for AI clients such as goose.
- Pre-configured lightweight subsets in
stacks/(e.g.,chat-only,voice-pipeline,rag-pipeline) needing as little as ~4.5 GB of RAM.
Documented constraints: 8 GB RAM minimum (16 GB recommended for 8B+ models), SpeakCrate, ParseCrate, and ScribeCrate Live are disabled by default, lightweight variants share default container names so only one should run at a time, and CUDA images are linux/amd64 only.
How engineers can use it
On a Linux host with Docker installed:
git clone https://github.com/hwdsl2/self-hosted-ai-stack
cd self-hosted-ai-stack
docker compose up -d
docker exec ollama ollama_manage --pull llama3.2:3b
./stack-check.sh
After startup, retrieve the GatewayCrate master key with docker exec litellm litellm_manage --showkey and the AnythingLLM admin password with docker exec anythingllm cat /app/server/storage/.initial_admin_password. Access the chat UI at http://<server-ip>:3001 and the LiteLLM Admin UI at http://<server-ip>:4000/ui.
For GPU acceleration, use docker compose -f docker-compose.cuda.yml up -d after installing the NVIDIA driver (575.57.08+ on Linux) and the NVIDIA Container Toolkit. The README includes worked curl examples for a voice pipeline (Whisper → LLM → Kokoro TTS), a RAG pipeline (embed → store in pgvector → query via GatewayCrate), and MCP tool calls hitting http://localhost:4000/v1/chat/completions. Podman users should install the podman-docker shim and use CDI for GPU passthrough, since Podman ignores the Compose deploy: GPU block.