Introduction to Model Deployment

Building a state-of-the-art machine learning model is only half the challenge. The other half β€” and often the harder half β€” is deploying that model to production where it serves real users, handles real traffic, and generates real value. Model deployment is the bridge between research and reality.

A deployed model must satisfy requirements that rarely matter in a notebook: low latency (inference in milliseconds), high throughput (serving thousands of requests per second), reliability (99.99% uptime), cost efficiency (minimizing GPU costs), and monitoring (catching performance degradation). A model that achieves 95% accuracy is worthless if it takes 10 seconds to respond or crashes under load.

This course covers everything needed to ship ML models to production. You'll learn about REST APIs with FastAPI, containerization with Docker, model optimization with ONNX and quantization, scaling with Kubernetes, and advanced serving frameworks like vLLM and Triton. By the end, you'll understand not just how to serve a single model, but how to build enterprise-grade ML infrastructure that scales reliably.

What You'll Learn

REST API Development

Build production-grade APIs with FastAPI that serve models efficiently, handle concurrent requests, and validate inputs. Learn request/response serialization and error handling.

Containerization & Docker

Package models, dependencies, and configurations into Docker containers that run identically everywhere β€” laptop, CI/CD pipeline, Kubernetes cluster.

Model Optimization

Compress and accelerate models using quantization (INT8, FP16), ONNX conversion, and distillation to reduce inference latency and cost.

Scale & Production

Deploy models at scale using Kubernetes, handle auto-scaling, implement monitoring and logging, and achieve sub-100ms latency under high load.

Prerequisites

Understanding of Python, basic knowledge of machine learning (training, inference, models), familiarity with command-line tools, and basic understanding of HTTP/REST APIs. Deep learning framework experience (PyTorch or TensorFlow) is helpful.

Why Model Deployment Matters

In industry, the cost of deploying a model often exceeds the cost of training it. A large language model might cost $10 million to train once, but if deployed inefficiently, could cost millions annually in inference GPU costs. Similarly, a 100ms latency increase can cost millions in lost user engagement. Model deployment isn't a footnote β€” it's central to an ML engineer's job.

Real-World Impact

Consider a recommendation system serving Spotify's 500M users:

Inference Latency Impact
50ms latency = 95% engagement
100ms Latency Cost
Lost 10% user engagement
Throughput Requirement
Serve 50,000 req/sec peak
Unoptimized Cost
$2M/month in GPU costs
Optimized Cost (Quantization)
$600K/month (70% savings)

Why deployment efficiency is business-critical

Key Challenges in Production

Latency

Models trained for accuracy must be optimized for speed. 50ms latency is acceptable; 500ms is not. Even 10ms improvements matter at scale.

Reliability

99.9% uptime means 9 hours of downtime per year are allowed. 99.99% (AWS/Google standard) means 44 minutes per year. Every crash is expensive.

Cost

GPU compute is expensive ($1-4 per hour per GPU). Optimize inference cost and throughput becomes a core business concern, not an afterthought.

Monitoring

Models degrade in production as data drifts. You need continuous monitoring to catch accuracy degradation, latency increases, and anomalous patterns.

Key Insight: A model that works perfectly in a Jupyter notebook but takes 2 seconds to respond in production is worthless. Deployment is not something you do after building a model β€” it's something you design for while building it.

Historical Evolution of Model Serving

Model deployment evolved alongside machine learning itself. Understanding this history helps you appreciate why different frameworks exist and which problems they solve.

2010s
Single Models
REST APIs + pickle files, manual scaling
2015
TensorFlow Serving
Google open-sources model serving, introduces batching
2017
Containerization
Docker adoption for reproducible deployments
2018
Kubernetes Era
Orchestrating containers at scale, auto-scaling
2020
LLM Serving
vLLM, Triton address unique challenges of large models
2023+
Multimodal Era
Serving LLMs, vision models, multimodal systems

The Evolution of Problems & Solutions

The Pickle Problem (2010s)

Early deployments pickled models and loaded them directly in Flask. This worked for small models on single servers but couldn't handle concurrent requests, batch processing, or multiple model versions. Every framework reinvented the wheel.

TensorFlow Serving (2015)

Google released TensorFlow Serving, the first production serving framework. It introduced crucial concepts: model versioning, batching requests for better GPU utilization, and separate model and serving code. This became the blueprint for modern serving systems.

Container Revolution (2017)

Docker made dependencies explicit and deployments reproducible. A model + Python version + CUDA version + dependencies could be packaged once and run identically everywhere. Combined with Kubernetes, this enabled truly scalable ML.

LLM Explosion (2023+)

GPT-3's 175B parameters broke existing serving frameworks. vLLM and Triton emerged to handle token generation efficiently using techniques like continuous batching and paged attention. The challenges of serving LLMs are fundamentally different from traditional models.

Core Concepts and Fundamentals

Model deployment introduces several new concepts that rarely appear in training. Understanding these fundamentals is crucial for building production systems.

Key Terms & Concepts

Inference vs Training

Training optimizes for accuracy using backpropagation and gradient updates. Inference runs a fixed model on new data to make predictions. Inference is what happens in production. A model might spend 10% of its cost during training and 90% during inference.

Latency

Time from receiving a request to returning a response. For real-time systems, sub-100ms is standard. Even 10ms improvements matter. Latency = model execution time + I/O + network overhead.

Throughput

Requests served per second. A model with 10ms latency can serve 100 requests per second per GPU (assuming optimal batching). Throughput is limited by memory bandwidth, not necessarily compute.

Batching

Instead of processing one request at a time, accumulate multiple requests and process them together. A model might do 10 inferences/sec on single samples but 1000 inferences/sec when processing batches of 100. The tradeoff is latency.

Model Quantization

Reduce model precision (FP32 β†’ INT8 or FP16). Cuts model size by 4x, speeds up inference 2-4x, but slightly reduces accuracy. Often accuracy loss is < 0.5% but cost savings are 70%+.

Cold Start & Warm Start

Cold start: first request pays penalty of loading model into GPU memory (100ms-1s). Warm start: subsequent requests reuse loaded model. Most production systems keep models loaded to avoid cold starts.

Model Versioning

Ability to run multiple versions of a model simultaneously (A/B testing), canary deployments (0.1% traffic on v2, 99.9% on v1), and safe rollbacks. Crucial for updates without downtime.

Deployment Architecture

Production ML systems follow a consistent architecture pattern, evolved through lessons from companies like Google, Meta, and Amazon.

The Standard Deployment Pipeline

1. Model Training & Validation
Train model, validate performance, export to standardized format (ONNX, SavedModel, PyTorch checkpoint).

2. Optimization & Quantization
Apply quantization, pruning, distillation to reduce latency and cost. Create optimized model variants for different hardware.

3. Containerization
Package model + serving code + dependencies into Docker image. Build for multiple architectures (x86, ARM, etc).

4. Local Testing
Run container locally, validate serving logic, benchmark latency. Test model loading, error handling, input validation.

5. Registry & Registry Management
Push image to container registry (Docker Hub, ECR, GCR). Tag with version, model name, hardware requirements.

6. Orchestration & Scaling
Deploy to Kubernetes. Configure auto-scaling: increase replicas if latency > 100ms or CPU > 80%.

7. Monitoring & Observability
Monitor latency (p50, p99), throughput, GPU utilization, accuracy drift. Alert on anomalies.

8. Continuous Deployment
New model versions roll out gradually: canary (1%), then 10%, then 50%, then 100%. Automatic rollback on errors.

Architecture Diagram

[Client] --HTTP--> [Load Balancer] ---> [Kubernetes Cluster]
                                                                           |
                                                                           +--> [Pod 1: Model + FastAPI]
                                                                           +--> [Pod 2: Model + FastAPI]
                                                                           +--> [Pod 3: Model + FastAPI]

[Monitoring: Prometheus] [Logging: ELK Stack] [Model Registry: MLflow]

Key Components of Production ML Systems

A production ML deployment isn't just the model β€” it's an ecosystem of components working together.

Essential Components

Model Serving Framework

Software that loads a model and handles HTTP requests. Examples: FastAPI (Python REST APIs), TensorFlow Serving (TF models), vLLM (LLMs), Triton (multi-framework). Each optimizes for different use cases.

Container Runtime

Docker or Podman packages the model, serving code, and dependencies. Enables reproducible deployments and portability across systems.

Orchestration Platform

Kubernetes manages containers at scale: scheduling, auto-scaling, rolling updates, secret management, persistent storage. Most enterprises use Kubernetes.

Model Registry

Central repository for trained models (e.g., MLflow, Hugging Face Hub, Model Zoo). Stores model artifacts, metadata, versioning, and lineage.

Monitoring & Observability

Prometheus scrapes metrics (latency, throughput, GPU util). Grafana visualizes them. ELK Stack or DataDog logs requests/errors. Alerts on anomalies.

Load Balancer

Routes requests across multiple serving instances. Kubernetes automatically provides this (Service resource). External load balancer (AWS ALB) distributes traffic to Kubernetes cluster.

CI/CD Pipeline

GitHub Actions, GitLab CI, or Jenkins automatically build Docker image, run tests, push to registry, deploy to staging, and promote to production.

Feature Store

System for managing input features (standardized, versioned, cached). Reduces training-serving skew by ensuring training and inference use identical features. Examples: Tecton, Feast.

Implementation: Building a Production-Ready API

Let's build a real, production-grade model serving API using FastAPI. This is how real production systems work.

Step 1: FastAPI Endpoint

Python β€” FastAPI Model Serving
from fastapi import FastAPI, HTTPException from pydantic import BaseModel import torch from transformers import pipeline import time from contextlib import asynccontextmanager # Global model cache _model_cache = {} @asynccontextmanager async def lifespan(app: FastAPI): # Startup: load model once print("Loading model...") _model_cache['classifier'] = pipeline( "sentiment-analysis", model="distilbert-base-uncased-finetuned-sst-2-english", device=0 if torch.cuda.is_available() else -1 ) print("Model loaded") yield # Shutdown: cleanup _model_cache.clear() app = FastAPI(lifespan=lifespan) class PredictionRequest(BaseModel): text: str class Config: examples = {"text": "This product is amazing!"} class PredictionResponse(BaseModel): label: str score: float latency_ms: float @app.post("/predict", response_model=PredictionResponse) async def predict(request: PredictionRequest): start = time.time() try: result = _model_cache['classifier'](request.text) latency_ms = (time.time() - start) * 1000 return PredictionResponse( label=result[0]['label'], score=round(result[0]['score'], 4), latency_ms=round(latency_ms, 2) ) except Exception as e: raise HTTPException(status_code=500, detail=str(e)) @app.get("/health") async def health_check(): return {"status": "healthy"}

Key production details: Model loads once at startup (expensive), then reused. Each request timed to monitor latency. Input validation via Pydantic. Error handling returns proper HTTP codes. Health check endpoint for load balancers.

Step 2: Docker Containerization

Python β€” Dockerfile for Model Serving
FROM nvidia/cuda:11.8.0-runtime-ubuntu22.04 WORKDIR /app # Install Python and dependencies RUN apt-get update && apt-get install -y python3.10 python3-pip # Copy requirements COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt # Copy app code COPY app.py . # Model will be downloaded at startup or mounted ENV TRANSFORMERS_CACHE=/models # Expose port EXPOSE 8000 # Health check HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \ CMD python -c "import requests; requests.get('http://localhost:8000/health')" # Run CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

Production features: Uses NVIDIA CUDA base image for GPU support. Caches model directory outside container. Health check for Kubernetes to detect crashed containers.

Step 3: Building and Running

Python β€” Build and Test
# Build image docker build -t model-server:v1.0 . # Run locally (test) docker run --gpus all -p 8000:8000 \ -e TRANSFORMERS_CACHE=/models \ model-server:v1.0 # Test endpoint curl -X POST http://localhost:8000/predict \ -H "Content-Type: application/json" \ -d '{"text": "This is amazing!"}' # Push to registry docker tag model-server:v1.0 myregistry.azurecr.io/model-server:v1.0 docker push myregistry.azurecr.io/model-server:v1.0

Advanced Optimization Techniques

Once basic serving works, production systems optimize for latency, cost, and reliability. These techniques can yield 10-100x improvements.

Model Quantization: INT8 Optimization

Python β€” INT8 Quantization with ONNX
import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer import onnxruntime as ort # Load original model model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased") tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased") # Export to ONNX dummy_input = tokenizer("test", return_tensors="pt") torch.onnx.export( model, tuple(dummy_input.values()), "model.onnx", input_names=["input_ids", "attention_mask"], output_names=["output"], opset_version=14 ) # Quantize to INT8 (4x smaller, 2-4x faster) from onnxruntime.quantization import quantize_dynamic, QuantType quantize_dynamic( "model.onnx", "model_quantized.onnx", weight_type=QuantType.QInt8, ) # Inference with ONNX sess = ort.InferenceSession("model_quantized.onnx") output = sess.run(None, {"input_ids": dummy_input['input_ids'].numpy()}) print(f"Output shape: {output[0].shape}") print(f"Accuracy loss: ~0.5% (typical for BERT quantization)")

vLLM: LLM-Specific Optimization

Python β€” vLLM for High-Throughput LLM Serving
from vllm import LLM, SamplingParams import time # Load LLM with vLLM (handles batching, paging, etc) llm = LLM( model="meta-llama/Llama-2-7b-hf", tensor_parallel_size=2, # Shard across 2 GPUs gpu_memory_utilization=0.9, # Use 90% GPU memory ) # vLLM implements continuous batching: # Instead of waiting for all requests in a batch, # processes requests as they arrive, yielding tokens incrementally sampling_params = SamplingParams(temperature=0.7, top_p=0.95, max_tokens=256) prompts = [ "What is machine learning?", "Explain quantum computing", "Write a Python function for..." ] # vLLM batches across GPU memory pages # Typical throughput: 500+ tokens/sec on 1 GPU (vs 10 on naive approach) outputs = llm.generate(prompts, sampling_params) for output in outputs: print(f"Prompt: {output.prompt}") print(f"Generated: {output.outputs[0].text}") print("---")

Triton Inference Server: Multi-Model Serving

Python β€” Triton Configuration for Multiple Models
# model_repository/ # β”œβ”€β”€ bert-model/ # β”‚ β”œβ”€β”€ config.pbtxt # β”‚ └── 1/ # β”‚ └── model.onnx # β”œβ”€β”€ gpt2-model/ # β”‚ β”œβ”€β”€ config.pbtxt # β”‚ └── 1/ # β”‚ └── model.onnx # └── ensemble-model/ # β”œβ”€β”€ config.pbtxt # └── 1/ # └── model.plan # bert-model/config.pbtxt: name: "bert-model" platform: "onnxruntime_onnx" max_batch_size: 128 input [ { name: "input_ids" data_type: TYPE_INT64 shape: [ 512 ] } ] output [ { name: "logits" data_type: TYPE_FP32 shape: [ 2 ] } ] # Run Triton server # docker run --gpus all -p 8000:8000 -p 8001:8001 \ # -v /path/to/model_repository:/models \ # nvcr.io/nvidia/tritonserver:23.04-py3 # Triton handles: # - Model batching (group requests for better GPU util) # - Model versioning (A/B test, canary deployment) # - Ensemble models (bert + gpt2 in sequence) # - Multiple backends (ONNX, TensorRT, TensorFlow, PyTorch) # - Dynamic batching (wait up to 5ms for more requests)

Performance Improvements Breakdown

FP32 Baseline
100ms latency
FP16 (Mixed Precision)
65ms (1.5x faster)
INT8 Quantization
40ms (2.5x faster)
INT8 + vLLM Batching
15ms (6.7x faster)
TensorRT (NVIDIA)
8ms (12x faster)

Latency improvements from various optimization techniques (cumulative)

Comparison of Serving Frameworks

Different frameworks optimize for different use cases. Choosing the right one matters for performance and cost.

Framework Best For Latency Throughput Learning Curve
FastAPI Quick prototypes, custom logic Good (50-200ms) Moderate (100-500 req/s) Easy
TensorFlow Serving TensorFlow models, batching Excellent (5-50ms) High (1000+ req/s) Medium
vLLM LLMs, token generation Good (50-500ms/token) Excellent (500+ tokens/s) Easy
Triton Multi-model, complex pipelines Excellent (5-100ms) Excellent (1000+ req/s) Hard
KServe Kubernetes-native, multi-framework Good (50-200ms) Moderate (100-500 req/s) Hard
Ray Serve Distributed serving, scalability Good (50-300ms) High (500+ req/s) Medium

Decision Tree

Simple REST API, small model? β†’ FastAPI

TensorFlow model, need high throughput? β†’ TensorFlow Serving

LLM with token generation? β†’ vLLM

Multiple models, complex pipelines? β†’ Triton

Kubernetes cluster with scale requirements? β†’ KServe or Ray Serve

Real-World Use Cases

Model deployment looks different across industries. Here's how leading companies serve ML in production.

Recommendation Systems (Netflix, Spotify)

Serve embeddings and ranking models to billions of users. Latency budget: 50ms. Solutions: Embedding cache (Redis), approximate nearest neighbors (Faiss), lightweight rankers. Netflix serves personalized recommendations in <30ms using GPU-accelerated approximate search.

Real-Time Translation (Google Translate)

Serve seq2seq models for 100+ language pairs with <500ms latency. Solutions: Model caching, quantization to FP16, batching requests. Google processes 100B+ translations per day using TensorFlow Serving with multiple model instances per GPU.

Autonomous Driving (Tesla, Waymo)

Serve vision models at 30+ fps on embedded hardware (not cloud). Latency budget: 33ms per frame. Solutions: Quantization to INT8, TensorRT optimization, multi-model inference. Waymo runs 20+ models per vehicle with <20ms total latency.

Conversational AI (ChatGPT, Claude)

Serve large language models with token generation. Latency budget: 20-50ms per token. Solutions: vLLM with paged attention, continuous batching, speculative decoding. OpenAI serves ChatGPT at 100+ tokens/sec per GPU using specialized inference optimization.

Fraud Detection (PayPal, Square)

Classify transactions in <100ms with high accuracy. Serve thousands of models (one per merchant). Solutions: Ensemble methods, lightweight models (linear/tree-based), feature caching. PayPal evaluates 1M+ transactions/sec using gradient boosted models.

Image Recognition (AWS, Google Cloud)

Serve vision models via REST API to millions of applications. Latency budget: 100-500ms. Solutions: Batching, model quantization, GPU pooling. AWS SageMaker serves 10B+ predictions monthly using TensorFlow/PyTorch inference containers.

Enterprise Deployment at Scale

Enterprise deployments differ from startups. They emphasize reliability, compliance, governance, and cost control.

Kubernetes Deployment with Auto-Scaling

Python β€” Kubernetes YAML for Auto-Scaling Deployment
apiVersion: apps/v1 kind: Deployment metadata: name: model-server spec: replicas: 3 selector: matchLabels: app: model-server template: metadata: labels: app: model-server spec: containers: - name: model-server image: myregistry.azurecr.io/model-server:v1.0 ports: - containerPort: 8000 resources: requests: memory: "4Gi" nvidia.com/gpu: "1" limits: memory: "6Gi" nvidia.com/gpu: "1" livenessProbe: httpGet: path: /health port: 8000 initialDelaySeconds: 30 periodSeconds: 10 readinessProbe: httpGet: path: /health port: 8000 initialDelaySeconds: 20 periodSeconds: 5 --- apiVersion: v1 kind: Service metadata: name: model-server-service spec: selector: app: model-server ports: - protocol: TCP port: 80 targetPort: 8000 type: LoadBalancer --- apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: model-server-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: model-server minReplicas: 3 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 - type: Resource resource: name: memory target: type: Utilization averageUtilization: 80

How it works: Kubernetes monitors CPU and memory. If CPU > 70%, adds a new pod. If requests exceed 3 pods' capacity (e.g., 50,000 β†’ 60,000 req/s), scales to 4 pods. If traffic drops, scales down. Liveness probe restarts crashed pods. Readiness probe removes unhealthy pods from load balancer.

Multi-Region Deployment for Global Availability

Architecture:

β€’ Global Load Balancer (AWS Route53, Google Cloud CDN) routes requests to nearest region
β€’ US East: Kubernetes cluster with 10 GPU nodes
β€’ EU West: Kubernetes cluster with 5 GPU nodes
β€’ APAC: Kubernetes cluster with 3 GPU nodes
β€’ Model Registry (central): Syncs model versions across regions
β€’ Monitoring (central): Collects metrics from all regions

Latency impact: Client in Singapore queries EU instead of US? 200ms latency improvement. Cost savings: Serve closer to users, reduce expensive inter-region traffic.

Cost Optimization in Production

Spot GPUs

Use cheaper spot instances for non-critical workloads. On AWS, spot V100 = 30% of on-demand cost. Tradeoff: 2-5 hour interruption risk.

Batching

Accumulate requests, process in batches. Throughput 10x higher (similar latency). Tradeoff: Increased latency for last request in batch.

Model Quantization

INT8/FP16 uses 4-8x less VRAM, enables smaller GPUs (T4 vs V100). Potential 1% accuracy loss, 70% cost reduction.

Multi-tenancy

Run multiple models on single GPU via shared memory. Requires isolation (memory limits), careful scheduling, but reduces idle GPU cost.

Common Mistakes in Model Deployment

Learning from others' mistakes can save months of debugging. Here are the most common deployment pitfalls.

Mistake 1: Not Benchmarking Latency in Advance

Building a model is exciting. Deploying it is urgent. But if latency requirements aren't clear before development, you might finish training a 2-second model that needs to serve in 100ms. Result: 6 months of optimization work. Solution: Define latency/throughput requirements before building the model.

Mistake 2: Training-Serving Skew

Model trained on preprocessed data, but serving code uses different preprocessing (different scaling, missing features, etc). Model makes poor predictions in production despite training accuracy of 95%. Solution: Use a feature store that both training and serving pipeline consume from.

Mistake 3: Not Testing Error Handling

Assuming models never crash. But what if input is malformed? Model out of memory? GPU fails? Production crashes silently or returns garbage. Solution: Input validation (Pydantic), exception handling, circuit breakers, fallback logic.

Mistake 4: Forgetting About Model Versioning

Deploy v2 of a model, traffic drops 30% (due to degradation), but can't rollback because no one documented v1. Solution: Always version models, keep previous versions, test v2 on 1% traffic before full rollout.

Mistake 5: Not Monitoring Data Drift

Model trained on 2023 data. In 2024, distribution shifts, accuracy drops to 60%. No one notices for 3 months. Solution: Monitor input distribution (via KL divergence), model predictions, actual labels (if available). Alert on drift.

Mistake 6: Deploying Bloated Models

Serving a 13B parameter model that could achieve 95% of the accuracy with 2B parameters (and be 6x faster). Solution: Early in development, test model size vs accuracy tradeoff. Distill if necessary.

Mistake 7: Ignoring Cold Start Time

Model takes 3 seconds to load into GPU memory. Every restart (deployments, crashes) has 3 second latency spike. With multiple pods restarting, cascades to 30 second downtime. Solution: Keep models warm, use persistent GPU memory, test startup time.

Mistake 8: Single Point of Failure

Only one instance of the model running. If it crashes, service is down until manual restart. Solution: Always run >= 3 replicas. Use auto-scaling. Have health checks.

Production Best Practices

Best practices are lessons learned by companies that got it wrong. Apply them from day one.

Development Phase

1. Define deployment requirements early: Latency (p50, p99), throughput (req/sec), accuracy, uptime SLA. Design model accordingly.

2. Build models for inference from start: Consider model size, quantizability, inference hardware. Avoid architectures that are hard to optimize.

3. Test with production-like inputs: Don't train on clean data, test on messy data. Simulate production traffic patterns.

4. Benchmark on target hardware: If deploying on T4 GPU, benchmark on T4 (not local GPU). Consider batch sizes, precision.

Deployment Phase

5. Use containers (Docker) for all deployments: Ensures reproducibility, simplifies dependency management, enables Kubernetes.

6. Version everything: Model artifacts, code, Docker images, configurations. Track lineage.

7. Implement health checks: Liveness (is the process alive?), readiness (can it handle traffic?), startup (how long to load model?)

8. Gradual rollout (canary deployment): Don't deploy v2 to 100% traffic at once. Deploy to 1%, then 5%, then 50%, with metrics comparison.

Production Phase

9. Monitor everything: Latency (p50, p99, p999), throughput, error rate, accuracy, resource utilization. Set up alerts.

10. Log for debugging: Log slow requests, errors, unusual inputs. Aggregrate with ELK or similar. But don't log raw predictions (privacy).

11. Implement circuit breakers: If model inference starts timing out, fail fast rather than cascading timeout.

12. Have a playbook for incidents: Model latency spikes, out of memory errors, accuracy degradation. Who to page? What to roll back?

Data & Governance

13. Monitor data drift: Track distribution of input features. Accuracy degradation often signals data shift.

14. Maintain audit trail: Log predictions for audit (legal requirement in some domains). Enable root cause analysis later.

15. Plan for retraining: Models degrade. Establish retraining cadence (weekly/monthly/quarterly) and automate it.

16. Document models: What data was it trained on? What performance guarantees? Known failure modes? Deprecation timeline?

Advanced Insights & State-of-the-Art

Beyond basics: cutting-edge techniques that leading companies use for competitive advantage.

Speculative Decoding for LLMs

Generating text token-by-token is slow. Speculative decoding uses a small fast model to generate draft tokens, then a large model verifies them in parallel. Result: Up to 2-3x faster text generation with no accuracy loss.

How it works: Small model generates 5 candidate tokens fast. Large model checks them all in parallel (much cheaper than generating them sequentially). If match, accept them all at once. If mismatch, use large model's choice. Net: 5 tokens generated in time of 1-2 large model calls.

Prefix Caching for Conversation

When chatting, you recompute attention for the same conversation history every turn. Prefix caching stores intermediate activations (KV cache) from previous turns. New query only computes attention on new tokens.

Performance gain: First turn of 1000-token conversation: 5 seconds. Second turn (new 50 tokens): 100ms (50x faster) instead of 5 seconds.

Model Merging & Ensemble

Train multiple models on different data or with different objectives, then merge them. A single merged model can match accuracy of ensemble while being faster.

Python β€” Simple Model Merging
# Train specialist models model_domain1 = train_model(data_domain1) # Finance domain model_domain2 = train_model(data_domain2) # Medical domain # Merge their weights (simple averaging or learned interpolation) merged_weights = 0.6 * model_domain1.weights + 0.4 * model_domain2.weights merged_model = create_model(merged_weights) # Result: One model that's good at both domains # No ensemble overhead (no need to run 2 models, just 1) # Faster inference, smaller deployment footprint

Continuous Batching vs Micro-Batching

Continuous Batching (vLLM): Instead of waiting for a full batch before processing, start processing as soon as requests arrive. Each request generates tokens independently. When one finishes, replace with waiting request. GPU rarely sits idle.

Micro-batching: Process requests in tiny batches (size 1-4) rapidly. Slightly higher latency than continuous batching but simpler to implement.

Unbatched (batch_size=1)
Throughput: 100 tokens/sec
Small Batching (batch_size=8)
Throughput: 600 tokens/sec
Continuous Batching
Throughput: 850 tokens/sec

Complete Code Examples

Real, runnable code for common deployment scenarios.

Example 1: End-to-End FastAPI + Docker

Python β€” app.py
from fastapi import FastAPI, HTTPException from fastapi.middleware.cors import CORSMiddleware from pydantic import BaseModel import torch from transformers import AutoTokenizer, AutoModelForSequenceClassification import logging import time logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) app = FastAPI(title="Sentiment API", version="1.0") # CORS for web clients app.add_middleware( CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"], ) model_cache = {} @app.on_event("startup") async def startup(): logger.info("Loading model...") tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english") model = AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english") if torch.cuda.is_available(): model.cuda() model.eval() model_cache['tokenizer'] = tokenizer model_cache['model'] = model logger.info("Model loaded successfully") class TextRequest(BaseModel): text: str class SentimentResponse(BaseModel): text: str sentiment: str confidence: float latency_ms: float @app.post("/predict", response_model=SentimentResponse) async def predict(request: TextRequest): if not request.text or len(request.text.strip()) == 0: raise HTTPException(status_code=400, detail="Text cannot be empty") start = time.time() tokenizer = model_cache['tokenizer'] model = model_cache['model'] try: inputs = tokenizer(request.text, return_tensors="pt", truncation=True, max_length=512) if torch.cuda.is_available(): inputs = {k: v.cuda() for k, v in inputs.items()} with torch.no_grad(): outputs = model(**inputs) logits = outputs.logits probs = torch.softmax(logits, dim=-1) pred_class = torch.argmax(probs, dim=-1).item() confidence = probs[0][pred_class].item() sentiment = "POSITIVE" if pred_class == 1 else "NEGATIVE" latency_ms = (time.time() - start) * 1000 return SentimentResponse( text=request.text[:100], sentiment=sentiment, confidence=round(confidence, 4), latency_ms=round(latency_ms, 2) ) except Exception as e: logger.error(f"Prediction error: {str(e)}") raise HTTPException(status_code=500, detail=str(e)) @app.get("/health") async def health(): return {"status": "healthy", "model": "loaded"} @app.get("/metrics") async def metrics(): return {"gpu_available": torch.cuda.is_available()}
Python β€” Dockerfile
FROM nvidia/cuda:11.8.0-runtime-ubuntu22.04 WORKDIR /app RUN apt-get update && apt-get install -y python3.10 python3-pip && rm -rf /var/lib/apt/lists/* COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY app.py . ENV TRANSFORMERS_CACHE=/models ENV PYTHONUNBUFFERED=1 EXPOSE 8000 HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3 \ CMD python3 -c "import requests; requests.get('http://localhost:8000/health').raise_for_status()" CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
Python β€” requirements.txt
fastapi==0.104.1 uvicorn[standard]==0.24.0 torch==2.0.1 transformers==4.34.0 pydantic==2.4.2 python-multipart==0.0.6

Example 2: ONNX Inference Optimization

Python β€” onnx_inference.py
import onnxruntime as ort import numpy as np from transformers import AutoTokenizer # Load ONNX model (pre-quantized to INT8) sess_options = ort.SessionOptions() sess_options.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL sess_options.execution_mode = ort.ExecutionMode.ORT_SEQUENTIAL session = ort.InferenceSession( "model_quantized.onnx", sess_options=sess_options, providers=['CUDAExecutionProvider', 'CPUExecutionProvider'] ) tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased") def predict_onnx(texts, batch_size=32): all_logits = [] for i in range(0, len(texts), batch_size): batch = texts[i:i+batch_size] # Tokenize inputs = tokenizer( batch, padding='max_length', truncation=True, max_length=128, return_tensors='np' ) # ONNX inference input_ids = inputs['input_ids'].astype(np.int64) attention_mask = inputs['attention_mask'].astype(np.int64) token_type_ids = inputs['token_type_ids'].astype(np.int64) output_names = session.get_outputs() output_names = [o.name for o in output_names] logits = session.run( output_names, { 'input_ids': input_ids, 'attention_mask': attention_mask, 'token_type_ids': token_type_ids, } ) all_logits.append(logits[0]) return np.vstack(all_logits) # Test texts = ["This is great!", "This is terrible."] * 100 predictions = predict_onnx(texts, batch_size=64) print(f"Predictions shape: {predictions.shape}") print(f"ONNX INT8 model: 4x smaller, 2-3x faster, <1% accuracy loss")

Example 3: vLLM Serving

Python β€” llm_serve.py
from vllm import LLM, SamplingParams from fastapi import FastAPI from pydantic import BaseModel app = FastAPI() # Load LLM with vLLM's continuous batching llm = LLM( model="meta-llama/Llama-2-7b-hf", tensor_parallel_size=1, gpu_memory_utilization=0.9, max_model_len=2048, ) class GenerateRequest(BaseModel): prompt: str max_tokens: int = 256 class GenerateResponse(BaseModel): prompt: str generated_text: str @app.post("/generate", response_model=GenerateResponse) async def generate(request: GenerateRequest): sampling_params = SamplingParams( temperature=0.7, top_p=0.95, max_tokens=request.max_tokens ) # vLLM handles continuous batching internally # Multiple concurrent requests are batched efficiently outputs = llm.generate([request.prompt], sampling_params) return GenerateResponse( prompt=request.prompt, generated_text=outputs[0].outputs[0].text ) # Run: uvicorn llm_serve:app --reload

Hands-On Exercises

Apply what you've learned with practical exercises.

Exercise 1: Build Your First FastAPI Endpoint

Create a simple FastAPI app that loads a pretrained HuggingFace model and serves predictions via /predict endpoint. Test with curl.

Python β€” Starter Code
from fastapi import FastAPI from pydantic import BaseModel app = FastAPI() class TextInput(BaseModel): text: str @app.post("/predict") async def predict(input: TextInput): # Load model and return predictions return {"prediction": "TODO"}

Exercise 2: Containerize the API

Write a Dockerfile for your FastAPI app. Build the image, test it with `docker run`, and push to a registry.

Python β€” Starter Code
FROM python:3.10 WORKDIR /app COPY requirements.txt . RUN pip install -r requirements.txt COPY app.py . CMD ["uvicorn", "app:app", "--host", "0.0.0.0"]

Exercise 3: Quantize and Benchmark

Load a model in FP32, quantize to INT8 using ONNX, benchmark latency of both. Calculate speedup and accuracy change.

Exercise 4: Deploy to Kubernetes

Write a Kubernetes deployment manifest for your containerized model server. Apply it to a cluster and test with `kubectl port-forward`.

Interview Questions

Common questions asked in ML engineering interviews about deployment.

1. How would you reduce latency of a model serving API from 500ms to 50ms? ▼
Answer: Multi-step optimization:

1. Profiling: Determine bottleneck. Is it model inference, I/O, preprocessing?
2. Model optimization: Quantization (FP32β†’INT8, 4x speedup). Distillation to smaller model. Pruning. Knowledge distillation.
3. Batching: Process requests in batches. 10x throughput (slight latency increase for last request).
4. Caching: Cache embeddings, feature preprocessing, model outputs if applicable.
5. Framework choice: vLLM for LLMs, Triton for complex pipelines. TensorFlow Serving for TF models.
6. Hardware: Use GPU. Faster GPU (V100 vs T4, 5x speedup). Multiple GPUs for ensemble.
7. Serving optimizations: Remove unnecessary preprocessing, minimize data copying, parallel model forward pass.

Typical result: 50-100x improvements combining several techniques.
2. How do you handle model versioning in production? ▼
Answer: Multi-strategy approach:

Version Storage:
β€’ Store model artifacts in model registry (MLflow, Hugging Face Hub) with version tags
β€’ Include metadata: training date, accuracy, model hyperparameters, dependencies

Deployment Strategy:
β€’ Never overwrite old model versions; always create new versions
β€’ Test new version on staging environment first
β€’ Canary deployment: Route 1% traffic to v2, monitor metrics, scale to 5%, 10%, 50%, 100%
β€’ Keep previous versions readily available for rollback

Serving Strategy:
β€’ Use model serving framework that supports multiple versions (Triton, TensorFlow Serving)
β€’ Implement A/B testing: Split traffic between v1 and v2, compare accuracy/latency
β€’ Automatic rollback if error rate increases >1% or latency >10%

Example: Deploy v1.5 to 1% traffic for 24 hours. If metrics are better, scale to 100%. If worse, automatic rollback to v1.4.
3. What's the difference between throughput and latency optimization? ▼
Answer: Different objectives, different approaches:

Latency Optimization (p99 < 100ms):
β€’ Minimize time per single request
β€’ Reduce model computation (quantization, pruning, smaller model)
β€’ Reduce I/O (pre-load model, local inference)
β€’ Trade throughput for latency (don't batch, or batch=1)
β€’ Use fastest GPU, CPU processor

Throughput Optimization (max req/sec):
β€’ Maximize requests per second
β€’ Batch requests aggressively (batch_size=128), slight latency increase
β€’ Parallel processing (multi-GPU, multi-model)
β€’ Can tolerate higher latency per request if total throughput increases

Example: Recommendation system needs latency. Spam detection can batch (can tolerate 100ms latency for throughput).
4. How would you set up monitoring for a production model? ▼
Answer: Multi-layer monitoring strategy:

Serving Metrics (Real-time):
β€’ Latency: p50, p99, p999 (not just mean)
β€’ Throughput: requests per second
β€’ Error rate: % of requests failing
β€’ Resource utilization: GPU memory, CPU, bandwidth

Model Metrics (Batch):
β€’ Accuracy: True positive rate, false positive rate on holdout test set
β€’ Data drift: Distribution of input features vs training distribution (KL divergence)
β€’ Prediction distribution: Are outputs changing unexpectedly?

Infrastructure Metrics:
β€’ Pod health: % replicas up, deployment success rate
β€’ Kubernetes health: Node availability, persistent storage

Tools:
β€’ Prometheus for metrics collection
β€’ Grafana for dashboards
β€’ ELK Stack for logs
β€’ Weights & Biases or similar for model-specific monitoring

Alerting:
β€’ Alert if p99 latency > 200ms
β€’ Alert if error rate > 1%
β€’ Alert if data drift detected
β€’ Alert if accuracy drops > 2%
5. How do you handle the cold start problem? ▼
Answer: Multi-strategy approach:

Problem: First request to a container incurs 1-5 second latency penalty for loading model into GPU memory. In auto-scaled systems with frequent deployment, cold starts are common and visible to users.

Solutions:

1. Warm Up on Startup:
   β€’ Send dummy requests in container startup script
   β€’ Load model and run inference before marking as ready
   β€’ Kubernetes readiness probe won't pass until warm

2. Model Preloading:
   β€’ Use Init Container in Kubernetes to load model from shared storage
   β€’ Pin model weights in GPU memory (don't evict)

3. Prevent Scale Down:
   β€’ Minimum replicas > 0 (keep at least 1 warm instance always running)
   β€’ Don't scale to zero, even during low traffic

4. Fast Startup:
   β€’ Use smaller models for initialization (distilled version)
   β€’ Cache compiled CUDA kernels

Tradeoff: Idle cost vs cold start latency. Usually worth keeping warm replica running.
6. What's the difference between batch inference and online serving? ▼
Answer: Different paradigms for different use cases:

Batch Inference:
β€’ Process large dataset (1M+ samples) at once
β€’ Latency not critical (can wait hours)
β€’ Throughput critical (maximize GPU utilization)
β€’ Example: Daily recommendation computation, overnight fraud scoring
β€’ Implementation: Spark, Kubernetes batch jobs, SageMaker Batch Transform

Online Serving:
β€’ Process single request in response to API call
β€’ Latency critical (must respond in 100ms)
β€’ Throughput secondary (though important at scale)
β€’ Example: Real-time search ranking, recommendation ranking
β€’ Implementation: FastAPI, Triton, vLLM

Mixed Approach (Lambdas Architecture):
β€’ Use batch inference for complex, heavy models overnight
β€’ Cache results (e.g., Redis)
β€’ Online serving looks up cached results
β€’ Combine batch + online for best of both worlds
7. How do you ensure feature consistency between training and serving? ▼
Answer: Feature stores and standardized pipelines:

The Problem (Training-Serving Skew):
β€’ Training pipeline: Complex feature engineering in Spark
β€’ Serving pipeline: Different code (maybe different language), slightly different logic
β€’ Result: Model sees different features at train time vs serve time
β€’ Impact: Accuracy drops from 95% β†’ 80% in production

Solution 1: Feature Store
β€’ Centralized system for feature definitions and computation
β€’ Both training and serving pipelines consume from feature store
β€’ Ensures identical features
β€’ Examples: Tecton, Feast, Vertex AI Feature Store

Solution 2: Unified Code
β€’ Write feature engineering in language-agnostic format (e.g., SQL)
β€’ Compile to both training and serving
β€’ Use same preprocessing for training and serving

Solution 3: Versioned Features
β€’ Version feature engineering code
β€’ When training model, record feature version used
β€’ At serving time, use same feature version

Solution 4: Testing
β€’ Compare features computed in training vs serving on same data
β€’ Alert if differences > threshold
8. Design a deployment system for serving 100 different models with varying latency requirements ▼
Answer: Complex multi-model orchestration:

Requirements Analysis:
β€’ 100 models, different sizes (1MB - 50GB)
β€’ Different latency needs (some need <50ms, some tolerate 1s)
β€’ Shared GPU resources
β€’ Cost optimization needed

Architecture:

1. Model Registry:
   β€’ MLflow: Store all 100 models with metadata (size, latency profile, accuracy)
   β€’ Tag each model with SLA requirements

2. Categorization by Latency:
   β€’ <50ms: Fast models (25 models, 2GB each) β†’ Triton on expensive GPUs (V100)
   β€’ 50-200ms: Medium models (50 models) β†’ vLLM or TF Serving on T4/A100
   β€’ >200ms: Slow models (25 models) β†’ Batch inference, Kubernetes jobs

3. Serving Strategy:
   β€’ API Gateway: Routes requests to appropriate serving infrastructure
   β€’ Model Router: Determines which serving engine for each model
   β€’ Load Balancer: Distributes to instances

4. GPU Sharing:
   β€’ Triton Ensemble: Run multiple small models on 1 GPU
   β€’ Memory isolation: Limit GPU memory per model
   β€’ Context switching: Switch between models based on traffic

5. Cost Optimization:
   β€’ Spot instances for batch inference (25% cost)
   β€’ Quantization for all models (reduce GPU requirements)
   β€’ Idle model unloading (free up GPU after N minutes)

6. Monitoring:
   β€’ Per-model metrics dashboard
   β€’ Alert if any model violates SLA
   β€’ Track GPU utilization per model

Frequently Asked Questions

Q1: Should I always use quantization? ▼
Quantization (FP32β†’INT8) typically reduces latency 2-4x with <1% accuracy loss. Always try it. The only exception: If your model is already quantized or if the accuracy loss exceeds your tolerance. For most production systems, quantization is free performance.
Q2: How do I choose between FastAPI, TensorFlow Serving, and Triton? ▼
FastAPI: Simple REST APIs, single models, quick prototyping. TensorFlow Serving: TensorFlow models needing batching and versioning. Triton: Multiple models, complex pipelines, need all the features. Choose the simplest tool that solves your problem.
Q3: Can I run multiple models on one GPU? ▼
Yes, but with caveats. Triton Ensemble can run multiple small models on one GPU using careful memory management. Tradeoff: Scheduling overhead, potential latency interference if models run simultaneously. Works best for small models or time-shared scheduling.
Q4: How often should I retrain models? ▼
Depends on data drift. Monitor input distribution: if drift detected, retrain. Typical cadence: daily for rapidly changing domains (news), weekly for stable domains (fraud detection), monthly for slowly-changing (recommendation). Automate retraining pipeline.
Q5: What's the cost difference between batch inference and online serving? ▼
Batch inference: $0.1/1000 predictions (efficient GPU usage). Online serving: $1-10/1000 predictions (GPU idle time, redundancy for availability). Use batch for non-time-critical tasks. Online for real-time requirements.
Q6: How do I handle model inference timeouts? ▼
Set a request timeout (e.g., 5 seconds). If inference takes longer, return an error or fallback result. Implement circuit breaker: if >10% requests timeout, stop sending traffic and alert on-call. Consider increasing timeout or optimizing model.
Q7: Should I version my code or just my models? ▼
Version both! Model without code context is useless (what preprocessing was used?). Use git for code, model registry for models. Each production model references specific code version.
Q8: Can I deploy on CPUs instead of GPUs? ▼
Yes, but 5-50x slower. For small models, CPUs sufficient. For real-time requirements, GPUs essential. Hybrid approach: Use CPUs for batch, GPUs for online serving.
Q9: How do I debug model serving issues in production? ▼
Check logs (ELK), monitor metrics (Prometheus/Grafana), profile model latency (time each component), test with sample data, compare to staging environment. Never modify code in production for debugging; create test deployment.
Q10: How do I handle privacy compliance (GDPR, etc) in deployments? ▼
Don't log raw predictions (models might leak PII). Log only: request ID, feature version, inference time, error. Store audit trail separately. Implement data retention: delete logs after N days. Test with legal team.

Summary & Key Takeaways

Model deployment is both art and science. Here are the essential principles:

Core Principles

1. Latency Matters: 100ms might not matter for batch inference, but it's critical for real-time systems. Measure it, optimize it, monitor it.

2. Reliability > Accuracy: A model that crashes is worthless. Invest in monitoring, alerts, and failover mechanisms before optimizing model accuracy.

3. Cost Discipline: GPU costs scale with traffic. Quantization, batching, and optimization save millions. Build cost optimization into model development.

4. Automation: Manual deployments are error-prone and slow. Invest in CI/CD pipelines. Models should deploy with same rigor as code.

5. Observability: You can't optimize what you don't measure. Instrument everything: latency, accuracy, resource usage, errors.

The Deployment Checklist

Before deploying to production, verify:

β˜‘ Latency/throughput requirements defined and met

β˜‘ Model quantized (INT8 or FP16) for production

β˜‘ Containerized (Docker) with health checks

β˜‘ Tested locally and in staging

β˜‘ Input validation and error handling implemented

β˜‘ Monitoring and alerting configured

β˜‘ Model versioning strategy in place

β˜‘ Rollback procedure documented

β˜‘ Canary deployment plan created

β˜‘ On-call runbook written

Next Steps

You now understand model deployment. To master it:

  1. Build: Deploy a simple model using FastAPI + Docker. Get comfortable with basics.
  2. Optimize: Quantize it, benchmark latency improvements, measure cost savings.
  3. Scale: Deploy to Kubernetes, implement auto-scaling, monitor in production.
  4. Advanced: Explore vLLM, Triton, multi-model serving, feature stores.
  5. Production: Lead a real model deployment at your company. Learn from mistakes.

Remember: The best model is the one serving users in production. Deployment isn't the final step β€” it's the beginning of the ML system's life. Invest accordingly.

Resources & Further Learning

Official Frameworks & Documentation

  • FastAPI: https://fastapi.tiangolo.com β€” Modern Python web framework
  • vLLM: https://docs.vllm.ai β€” LLM serving with continuous batching
  • TensorFlow Serving: https://www.tensorflow.org/tfx/guide/serving β€” Production TF model serving
  • Triton Inference Server: https://docs.nvidia.com/triton β€” Multi-model multi-framework serving
  • Kubernetes: https://kubernetes.io/docs β€” Container orchestration
  • Docker: https://docs.docker.com β€” Containerization platform
  • MLflow: https://mlflow.org β€” Model registry and experiment tracking

Recommended Books & Papers

Courses & Tutorials

Tools & Platforms

Community & Blogs

  • Chip Huyen's Blog: huyenchip.com β€” ML systems and deployment insights
  • vLLM Team Blogs: Latest LLM serving techniques
  • NVIDIA Developer Blog: GPU optimization and TensorRT
  • Kubernetes Community: kubernetes.io/community
  • Fast.ai Forums: Active community discussion on deployment