Mykhailo Pavlov
Back to Projects

Remote LLM Server

Released
Released Mar 2026
DockerOllamaLLMGPUAPI
Summary

A one-command Dockerized Ollama server with GPU acceleration that serves private LLMs to every device on your network.

The Problem

When building Anagnosi, I needed a reliable way to run large language models locally and serve them to multiple clients without duplicating model storage on each machine.

The Solution

Built a Dockerized Ollama server with GPU passthrough via NVIDIA Container Toolkit. Deploys with a single docker-compose command. Exposes a REST API on port 11434, accessible from any device on the local network.

Key features:

  • Runs Ollama inside Docker with full GPU acceleration
  • Persists downloaded models in a named Docker volume so they survive container restarts
  • Keeps models loaded in memory for 24 hours to eliminate cold-start latency
  • Falls back to CPU automatically if no GPU is available