text-generation-inference
Text-Generation-Inference Overview
Secure your stack with a hardened Text-Generation-Inference image freshly-built by Minimus. Minimus images always include the most up-to-date package version for all packages and dependencies contained in the image.
Text Generation Inference (TGI) lets teams serve open-source large language models from the Hugging Face Hub in production without building a custom serving stack. It handles the performance-critical work of serving models — request batching, GPU memory management, and quantization — so a wide range of model architectures can run with low latency and high throughput on a shared set of GPUs. TGI also exposes an OpenAI-compatible Messages API, so it can act as a drop-in backend for existing chat and completion clients.
Try It Out
The Minimus Text-Generation-Inference image is designed to run on a machine equipped with an NVIDIA GPU. This guide demonstrates how to serve a model and query it.
Start the Server
Run the container with GPU access enabled, serving the Qwen2.5-1.5B-Instruct model:
docker run --rm -d --gpus all \
--shm-size 1g \
-v ~/.cache/huggingface:/data \
-p 8080:80 \
reg.mini.dev/text-generation-inference \
--model-id Qwen/Qwen2.5-1.5B-InstructThe server will start listening on http://localhost:8080. Model download and loading may take a few minutes on first run.
To serve a gated or private model, pass your Hugging Face token:
docker run --rm -d --gpus all \
--shm-size 1g \
-v ~/.cache/huggingface:/data \
-p 8080:80 \
-e HF_TOKEN=<your-Hugging-Face-token> \
reg.mini.dev/text-generation-inference \
--model-id meta-llama/Llama-3.1-8B-InstructSend a Test Request
From another terminal, send a chat completion request using TGI's OpenAI-compatible Messages API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello, what can you do?"}
]
}'You can also query TGI's native /generate endpoint directly:
curl http://localhost:8080/generate \
-H "Content-Type: application/json" \
-d '{
"inputs": "Hello, what can you do?",
"parameters": {"max_new_tokens": 50}
}'For more deployment options (multi-GPU sharding, Kubernetes, quantization), see the official TGI documentation.
Technical Considerations
The Text-Generation-Inference image provided by Minimus is a slim, security-hardened alternative to the public image from GitHub Container Registry. The images are largely interchangeable, with a few differences as noted below.
Text-Generation-Inference built by Minimus:
- The Minimus Text-Generation-Inference image is designed to run on a machine equipped with an NVIDIA GPU.
- Runs as root to support required functions in alignment with the public image.
- Drill down on the version specification tab to see the default user, listening ports, entrypoint, volumes, environment variables, etc.
The Payoff
A hardened, minimal image that will remain more secure for the long run and accrue vulnerabilities at a slower rate.
- See the risk reduction dashboard for a detailed CVE comparison over the past 30 days.
- Review the compliance report to see the default hardening and security configurations for the image.
Terms & Info
Trademark
This catalog is published by Minimus. All product names, logos, and marks, other than those belonging to Minimus, shown are owned by their respective rights holders and appear here only to identify the open source software each image contains. Minimus claims no ownership of those marks and implies no affiliation with, endorsement by, certification by, or sponsorship by any rights holder.
Disclaimer
Images are provided "as-is" without warranty of any kind. "Hardened" refers to the security configuration applied at the time of build and does not constitute a guarantee of ongoing security or absence of vulnerabilities. The free tier is provided without support, SLA, or guaranteed patching timelines. Security updates may be applied to paid subscriptions before or instead of free tier images. By pulling or using any image you agree to our Terms of Use.