Back to Projects
Remote LLM Server
Released
Released Mar 2026
DockerOllamaLLMGPUAPI
Summary
A one-command Dockerized Ollama server with GPU acceleration that serves private LLMs to every device on your network.
The Problem
When building Anagnosi, I needed a reliable way to run large language models locally and serve them to multiple clients without duplicating model storage on each machine.
The Solution
Built a Dockerized Ollama server with GPU passthrough via NVIDIA Container Toolkit. Deploys with a single docker-compose command. Exposes a REST API on port 11434, accessible from any device on the local network.
Key features:
- Runs Ollama inside Docker with full GPU acceleration
- Persists downloaded models in a named Docker volume so they survive container restarts
- Keeps models loaded in memory for 24 hours to eliminate cold-start latency
- Falls back to CPU automatically if no GPU is available