Model Deployment Guide¶
This guide explains how to save, deploy, and manage models across multiple self-hosted Ollama instances.
Overview¶
When working with self-hosted Ollama services, you often need to: - Deploy custom models to production servers - Share models between development and staging environments - Backup models for disaster recovery - Transfer models between team members
This project includes scripts and make commands to streamline these workflows.
Quick Start¶
Save a Model¶
Save a model's Modelfile for deployment:
# Interactive
make save-model
# Or specify directly
bash scripts/save-model.sh my-chatbot
# Save to custom location
bash scripts/save-model.sh my-chatbot ./backups/production
This creates a Modelfile at ./models/saved/my-chatbot.Modelfile (or your specified location).
Deploy a Model¶
Deploy a saved Modelfile to the current Ollama instance:
# Interactive
make deploy-model
# Or specify directly
bash scripts/deploy-model.sh ./models/saved/my-chatbot.Modelfile
# Deploy with a different name
bash scripts/deploy-model.sh ./models/saved/my-chatbot.Modelfile production-chatbot
Backup All Models¶
Create a timestamped backup of all custom models:
# Default location: ./backups/models/YYYYMMDD_HHMMSS/
make backup-models
# Or specify custom location
bash scripts/backup-models.sh /mnt/backup/ollama-models
Full Export/Import (Weights Included)¶
make save-model exports only the recipe (Modelfile): the target instance must
re-download the base model from the Ollama library. When the target is air-gapped
or on a slow connection, export the model with its weights instead:
# Interactive
make export-full
# Or specify directly (default output: ./backups/full/)
bash scripts/export-model-full.sh my-chatbot
bash scripts/export-model-full.sh llama3.2:1b ./exports
This creates a tar archive containing the model's manifest and all its blobs (weights, system prompt, parameters). Transfer it to the target machine, then:
# Interactive (lists archives in ./backups/full/)
make import-full
# Or specify directly
bash scripts/import-model-full.sh ./backups/full/my-chatbot-20260704_120000.tar
The model is available immediately after import — nothing is downloaded.
Notes:
- Archives include the full weights: expect 1-5 GB+ per model. For everyday
sharing between connected instances, make save-model stays the better option.
- The archive layout is Ollama's internal storage format. Keep source and target
on similar Ollama versions — pin OLLAMA_IMAGE_TAG in .env on both sides.
- Blobs are content-addressed: importing a model whose base blobs already exist
on the target simply overwrites them (no duplication).
Deployment Workflows¶
Workflow 1: Development to Production¶
-
On development server, create and test your model:
# Create custom model bash scripts/create-custom-model.sh my-app-assistant ./models/custom/app-assistant/Modelfile # Test it docker compose exec ollama ollama run my-app-assistant "Test prompt" -
Save the model:
bash scripts/save-model.sh my-app-assistant # Creates: ./models/saved/my-app-assistant.Modelfile -
Transfer to production server:
scp ./models/saved/my-app-assistant.Modelfile user@prod-server:/opt/ollama/models/saved/ -
On production server, deploy:
cd /opt/ollama bash scripts/deploy-model.sh ./models/saved/my-app-assistant.Modelfile -
Verify deployment:
docker compose exec ollama ollama list docker compose exec ollama ollama run my-app-assistant "Test prompt"
Workflow 2: Team Collaboration¶
Share models via version control:
-
Developer A creates a model:
bash scripts/create-custom-model.sh team-assistant ./models/custom/team/Modelfile bash scripts/save-model.sh team-assistant ./models/custom/team/ -
Commit to repository:
git add models/custom/team/ git commit -m "Add team assistant model" git push -
Developer B pulls and deploys:
git pull bash scripts/deploy-model.sh ./models/custom/team/team-assistant.Modelfile
Workflow 3: Disaster Recovery¶
Regular backups ensure you can recover from data loss:
-
Schedule regular backups (e.g., daily cron job):
# Add to crontab 0 2 * * * cd /opt/ollama && bash scripts/backup-models.sh /mnt/backup/ollama -
If disaster strikes, restore from backup:
# List available backups ls -la /mnt/backup/ollama/ # Deploy models from specific backup for modelfile in /mnt/backup/ollama/20240315_020000/*.Modelfile; do bash scripts/deploy-model.sh "$modelfile" done
Workflow 4: Multi-Environment Deployment¶
Deploy the same model to dev, staging, and prod:
-
Create master Modelfile in version control:
# models/production/customer-support/Modelfile FROM llama3.2:3b PARAMETER temperature 0.6 SYSTEM """You are a helpful customer support assistant...""" -
Deploy to all environments:
# Development ssh dev-server "cd /opt/ollama && bash scripts/deploy-model.sh ./models/production/customer-support/Modelfile" # Staging ssh stage-server "cd /opt/ollama && bash scripts/deploy-model.sh ./models/production/customer-support/Modelfile" # Production ssh prod-server "cd /opt/ollama && bash scripts/deploy-model.sh ./models/production/customer-support/Modelfile"
What Gets Saved/Deployed?¶
Saved in Modelfile¶
- Base model reference (FROM)
- All parameters (temperature, num_ctx, etc.)
- System prompt (SYSTEM)
- Custom template (TEMPLATE, if defined)
- Few-shot examples (MESSAGE, if defined)
- Adapter references (ADAPTER, if defined)
NOT Saved in Modelfile¶
- Base model weights: The underlying model (e.g., llama3.2:3b) must be available on the target instance
- GGUF files: External model files must be copied separately
- LoRA adapters: Adapter files must be transferred separately
Important Notes¶
Base Model Availability¶
When you deploy a Modelfile, the target instance must have access to the base model:
FROM llama3.2:3b # This model must exist on target instance
Before deploying, ensure base model is available:
# On target instance
docker compose exec ollama ollama pull llama3.2:3b
External GGUF Models¶
If your model uses a local GGUF file:
FROM /data/gguf/my-custom-model.gguf
You must transfer the GGUF file separately:
scp ./data/gguf/my-custom-model.gguf user@target:/opt/ollama/data/gguf/
LoRA Adapters¶
If your model uses adapters:
FROM llama3.2:3b
ADAPTER /data/adapters/my-adapter.bin
Transfer the adapter file:
scp ./data/adapters/my-adapter.bin user@target:/opt/ollama/data/adapters/
Automation Scripts¶
Automated Deployment Script Example¶
Create a deployment automation script:
#!/bin/bash
# deploy-to-production.sh
MODEL_NAME=$1
PROD_SERVER="user@prod-server"
PROD_PATH="/opt/ollama"
# Save model
bash scripts/save-model.sh "$MODEL_NAME"
# Transfer
scp "./models/saved/${MODEL_NAME}.Modelfile" "$PROD_SERVER:$PROD_PATH/models/saved/"
# Deploy remotely
ssh "$PROD_SERVER" "cd $PROD_PATH && bash scripts/deploy-model.sh ./models/saved/${MODEL_NAME}.Modelfile"
echo "✅ Deployed $MODEL_NAME to production"
Automated Backup Script Example¶
#!/bin/bash
# backup-to-s3.sh
BACKUP_DIR="/tmp/ollama-backup"
# Create backup
bash scripts/backup-models.sh "$BACKUP_DIR"
# Upload to S3
LATEST_BACKUP=$(ls -t "$BACKUP_DIR" | head -1)
aws s3 sync "$BACKUP_DIR/$LATEST_BACKUP" "s3://my-bucket/ollama-backups/$LATEST_BACKUP/"
# Cleanup local backup
rm -rf "$BACKUP_DIR"
echo "✅ Backup uploaded to S3"
Troubleshooting¶
Model Not Found After Deployment¶
Issue: Deployed model doesn't appear in ollama list
Solutions:
1. Check base model exists: docker compose exec ollama ollama pull <base-model>
2. Check Modelfile syntax: Look for errors in deployment output
3. Verify container has access to referenced files (GGUF, adapters)
Different Behavior on Target Instance¶
Issue: Model behaves differently on target vs source
Possible causes: 1. Different base model versions 2. Missing adapter files 3. Hardware differences (CPU vs GPU)
Solution: Ensure identical base models and all dependencies are present
Backup Script Skips Models¶
Issue: Some models aren't included in backups
Reason: Script only exports custom models, not base models from Ollama library
Solution: This is intended behavior. Base models should be pulled from Ollama library on target instances
Best Practices¶
- Version Control: Keep Modelfiles in Git for tracking changes
- Naming Convention: Use descriptive names (e.g.,
customer-support-v2,code-assistant-python) - Regular Backups: Schedule automated backups for custom models
- Test Before Production: Always test deployments in staging first
- Document Dependencies: Note base models and adapters in README or documentation
- Environment Parity: Use same base model versions across environments