Introduction to Model Deployment
Building a state-of-the-art machine learning model is only half the challenge. The other half β and often the harder half β is deploying that model to production where it serves real users, handles real traffic, and generates real value. Model deployment is the bridge between research and reality.
A deployed model must satisfy requirements that rarely matter in a notebook: low latency (inference in milliseconds), high throughput (serving thousands of requests per second), reliability (99.99% uptime), cost efficiency (minimizing GPU costs), and monitoring (catching performance degradation). A model that achieves 95% accuracy is worthless if it takes 10 seconds to respond or crashes under load.
This course covers everything needed to ship ML models to production. You'll learn about REST APIs with FastAPI, containerization with Docker, model optimization with ONNX and quantization, scaling with Kubernetes, and advanced serving frameworks like vLLM and Triton. By the end, you'll understand not just how to serve a single model, but how to build enterprise-grade ML infrastructure that scales reliably.
What You'll Learn
REST API Development
Build production-grade APIs with FastAPI that serve models efficiently, handle concurrent requests, and validate inputs. Learn request/response serialization and error handling.
Containerization & Docker
Package models, dependencies, and configurations into Docker containers that run identically everywhere β laptop, CI/CD pipeline, Kubernetes cluster.
Model Optimization
Compress and accelerate models using quantization (INT8, FP16), ONNX conversion, and distillation to reduce inference latency and cost.
Scale & Production
Deploy models at scale using Kubernetes, handle auto-scaling, implement monitoring and logging, and achieve sub-100ms latency under high load.
Prerequisites
Understanding of Python, basic knowledge of machine learning (training, inference, models), familiarity with command-line tools, and basic understanding of HTTP/REST APIs. Deep learning framework experience (PyTorch or TensorFlow) is helpful.
Why Model Deployment Matters
In industry, the cost of deploying a model often exceeds the cost of training it. A large language model might cost $10 million to train once, but if deployed inefficiently, could cost millions annually in inference GPU costs. Similarly, a 100ms latency increase can cost millions in lost user engagement. Model deployment isn't a footnote β it's central to an ML engineer's job.
Real-World Impact
Consider a recommendation system serving Spotify's 500M users:
Why deployment efficiency is business-critical
Key Challenges in Production
Latency
Models trained for accuracy must be optimized for speed. 50ms latency is acceptable; 500ms is not. Even 10ms improvements matter at scale.
Reliability
99.9% uptime means 9 hours of downtime per year are allowed. 99.99% (AWS/Google standard) means 44 minutes per year. Every crash is expensive.
Cost
GPU compute is expensive ($1-4 per hour per GPU). Optimize inference cost and throughput becomes a core business concern, not an afterthought.
Monitoring
Models degrade in production as data drifts. You need continuous monitoring to catch accuracy degradation, latency increases, and anomalous patterns.
Key Insight: A model that works perfectly in a Jupyter notebook but takes 2 seconds to respond in production is worthless. Deployment is not something you do after building a model β it's something you design for while building it.
Historical Evolution of Model Serving
Model deployment evolved alongside machine learning itself. Understanding this history helps you appreciate why different frameworks exist and which problems they solve.
Single Models
REST APIs + pickle files, manual scaling
TensorFlow Serving
Google open-sources model serving, introduces batching
Containerization
Docker adoption for reproducible deployments
Kubernetes Era
Orchestrating containers at scale, auto-scaling
LLM Serving
vLLM, Triton address unique challenges of large models
Multimodal Era
Serving LLMs, vision models, multimodal systems
The Evolution of Problems & Solutions
The Pickle Problem (2010s)
Early deployments pickled models and loaded them directly in Flask. This worked for small models on single servers but couldn't handle concurrent requests, batch processing, or multiple model versions. Every framework reinvented the wheel.
TensorFlow Serving (2015)
Google released TensorFlow Serving, the first production serving framework. It introduced crucial concepts: model versioning, batching requests for better GPU utilization, and separate model and serving code. This became the blueprint for modern serving systems.
Container Revolution (2017)
Docker made dependencies explicit and deployments reproducible. A model + Python version + CUDA version + dependencies could be packaged once and run identically everywhere. Combined with Kubernetes, this enabled truly scalable ML.
LLM Explosion (2023+)
GPT-3's 175B parameters broke existing serving frameworks. vLLM and Triton emerged to handle token generation efficiently using techniques like continuous batching and paged attention. The challenges of serving LLMs are fundamentally different from traditional models.
Core Concepts and Fundamentals
Model deployment introduces several new concepts that rarely appear in training. Understanding these fundamentals is crucial for building production systems.
Key Terms & Concepts
Inference vs Training
Training optimizes for accuracy using backpropagation and gradient updates. Inference runs a fixed model on new data to make predictions. Inference is what happens in production. A model might spend 10% of its cost during training and 90% during inference.
Latency
Time from receiving a request to returning a response. For real-time systems, sub-100ms is standard. Even 10ms improvements matter. Latency = model execution time + I/O + network overhead.
Throughput
Requests served per second. A model with 10ms latency can serve 100 requests per second per GPU (assuming optimal batching). Throughput is limited by memory bandwidth, not necessarily compute.
Batching
Instead of processing one request at a time, accumulate multiple requests and process them together. A model might do 10 inferences/sec on single samples but 1000 inferences/sec when processing batches of 100. The tradeoff is latency.
Model Quantization
Reduce model precision (FP32 β INT8 or FP16). Cuts model size by 4x, speeds up inference 2-4x, but slightly reduces accuracy. Often accuracy loss is < 0.5% but cost savings are 70%+.
Cold Start & Warm Start
Cold start: first request pays penalty of loading model into GPU memory (100ms-1s). Warm start: subsequent requests reuse loaded model. Most production systems keep models loaded to avoid cold starts.
Model Versioning
Ability to run multiple versions of a model simultaneously (A/B testing), canary deployments (0.1% traffic on v2, 99.9% on v1), and safe rollbacks. Crucial for updates without downtime.
Deployment Architecture
Production ML systems follow a consistent architecture pattern, evolved through lessons from companies like Google, Meta, and Amazon.
The Standard Deployment Pipeline
Train model, validate performance, export to standardized format (ONNX, SavedModel, PyTorch checkpoint).
2. Optimization & Quantization
Apply quantization, pruning, distillation to reduce latency and cost. Create optimized model variants for different hardware.
3. Containerization
Package model + serving code + dependencies into Docker image. Build for multiple architectures (x86, ARM, etc).
4. Local Testing
Run container locally, validate serving logic, benchmark latency. Test model loading, error handling, input validation.
5. Registry & Registry Management
Push image to container registry (Docker Hub, ECR, GCR). Tag with version, model name, hardware requirements.
6. Orchestration & Scaling
Deploy to Kubernetes. Configure auto-scaling: increase replicas if latency > 100ms or CPU > 80%.
7. Monitoring & Observability
Monitor latency (p50, p99), throughput, GPU utilization, accuracy drift. Alert on anomalies.
8. Continuous Deployment
New model versions roll out gradually: canary (1%), then 10%, then 50%, then 100%. Automatic rollback on errors.
Architecture Diagram
[Client] --HTTP--> [Load Balancer] ---> [Kubernetes Cluster]
|
+--> [Pod 1: Model + FastAPI]
+--> [Pod 2: Model + FastAPI]
+--> [Pod 3: Model + FastAPI]
[Monitoring: Prometheus] [Logging: ELK Stack] [Model Registry: MLflow]
Key Components of Production ML Systems
A production ML deployment isn't just the model β it's an ecosystem of components working together.
Essential Components
Model Serving Framework
Software that loads a model and handles HTTP requests. Examples: FastAPI (Python REST APIs), TensorFlow Serving (TF models), vLLM (LLMs), Triton (multi-framework). Each optimizes for different use cases.
Container Runtime
Docker or Podman packages the model, serving code, and dependencies. Enables reproducible deployments and portability across systems.
Orchestration Platform
Kubernetes manages containers at scale: scheduling, auto-scaling, rolling updates, secret management, persistent storage. Most enterprises use Kubernetes.
Model Registry
Central repository for trained models (e.g., MLflow, Hugging Face Hub, Model Zoo). Stores model artifacts, metadata, versioning, and lineage.
Monitoring & Observability
Prometheus scrapes metrics (latency, throughput, GPU util). Grafana visualizes them. ELK Stack or DataDog logs requests/errors. Alerts on anomalies.
Load Balancer
Routes requests across multiple serving instances. Kubernetes automatically provides this (Service resource). External load balancer (AWS ALB) distributes traffic to Kubernetes cluster.
CI/CD Pipeline
GitHub Actions, GitLab CI, or Jenkins automatically build Docker image, run tests, push to registry, deploy to staging, and promote to production.
Feature Store
System for managing input features (standardized, versioned, cached). Reduces training-serving skew by ensuring training and inference use identical features. Examples: Tecton, Feast.
Implementation: Building a Production-Ready API
Let's build a real, production-grade model serving API using FastAPI. This is how real production systems work.
Step 1: FastAPI Endpoint
Key production details: Model loads once at startup (expensive), then reused. Each request timed to monitor latency. Input validation via Pydantic. Error handling returns proper HTTP codes. Health check endpoint for load balancers.
Step 2: Docker Containerization
Production features: Uses NVIDIA CUDA base image for GPU support. Caches model directory outside container. Health check for Kubernetes to detect crashed containers.
Step 3: Building and Running
Advanced Optimization Techniques
Once basic serving works, production systems optimize for latency, cost, and reliability. These techniques can yield 10-100x improvements.
Model Quantization: INT8 Optimization
vLLM: LLM-Specific Optimization
Triton Inference Server: Multi-Model Serving
Performance Improvements Breakdown
Latency improvements from various optimization techniques (cumulative)
Comparison of Serving Frameworks
Different frameworks optimize for different use cases. Choosing the right one matters for performance and cost.
Decision Tree
Simple REST API, small model? β FastAPI
TensorFlow model, need high throughput? β TensorFlow Serving
LLM with token generation? β vLLM
Multiple models, complex pipelines? β Triton
Kubernetes cluster with scale requirements? β KServe or Ray Serve
Real-World Use Cases
Model deployment looks different across industries. Here's how leading companies serve ML in production.
Recommendation Systems (Netflix, Spotify)
Serve embeddings and ranking models to billions of users. Latency budget: 50ms. Solutions: Embedding cache (Redis), approximate nearest neighbors (Faiss), lightweight rankers. Netflix serves personalized recommendations in <30ms using GPU-accelerated approximate search.
Real-Time Translation (Google Translate)
Serve seq2seq models for 100+ language pairs with <500ms latency. Solutions: Model caching, quantization to FP16, batching requests. Google processes 100B+ translations per day using TensorFlow Serving with multiple model instances per GPU.
Autonomous Driving (Tesla, Waymo)
Serve vision models at 30+ fps on embedded hardware (not cloud). Latency budget: 33ms per frame. Solutions: Quantization to INT8, TensorRT optimization, multi-model inference. Waymo runs 20+ models per vehicle with <20ms total latency.
Conversational AI (ChatGPT, Claude)
Serve large language models with token generation. Latency budget: 20-50ms per token. Solutions: vLLM with paged attention, continuous batching, speculative decoding. OpenAI serves ChatGPT at 100+ tokens/sec per GPU using specialized inference optimization.
Fraud Detection (PayPal, Square)
Classify transactions in <100ms with high accuracy. Serve thousands of models (one per merchant). Solutions: Ensemble methods, lightweight models (linear/tree-based), feature caching. PayPal evaluates 1M+ transactions/sec using gradient boosted models.
Image Recognition (AWS, Google Cloud)
Serve vision models via REST API to millions of applications. Latency budget: 100-500ms. Solutions: Batching, model quantization, GPU pooling. AWS SageMaker serves 10B+ predictions monthly using TensorFlow/PyTorch inference containers.
Enterprise Deployment at Scale
Enterprise deployments differ from startups. They emphasize reliability, compliance, governance, and cost control.
Kubernetes Deployment with Auto-Scaling
How it works: Kubernetes monitors CPU and memory. If CPU > 70%, adds a new pod. If requests exceed 3 pods' capacity (e.g., 50,000 β 60,000 req/s), scales to 4 pods. If traffic drops, scales down. Liveness probe restarts crashed pods. Readiness probe removes unhealthy pods from load balancer.
Multi-Region Deployment for Global Availability
β’ Global Load Balancer (AWS Route53, Google Cloud CDN) routes requests to nearest region
β’ US East: Kubernetes cluster with 10 GPU nodes
β’ EU West: Kubernetes cluster with 5 GPU nodes
β’ APAC: Kubernetes cluster with 3 GPU nodes
β’ Model Registry (central): Syncs model versions across regions
β’ Monitoring (central): Collects metrics from all regions
Latency impact: Client in Singapore queries EU instead of US? 200ms latency improvement. Cost savings: Serve closer to users, reduce expensive inter-region traffic.
Cost Optimization in Production
Spot GPUs
Use cheaper spot instances for non-critical workloads. On AWS, spot V100 = 30% of on-demand cost. Tradeoff: 2-5 hour interruption risk.
Batching
Accumulate requests, process in batches. Throughput 10x higher (similar latency). Tradeoff: Increased latency for last request in batch.
Model Quantization
INT8/FP16 uses 4-8x less VRAM, enables smaller GPUs (T4 vs V100). Potential 1% accuracy loss, 70% cost reduction.
Multi-tenancy
Run multiple models on single GPU via shared memory. Requires isolation (memory limits), careful scheduling, but reduces idle GPU cost.
Common Mistakes in Model Deployment
Learning from others' mistakes can save months of debugging. Here are the most common deployment pitfalls.
Mistake 1: Not Benchmarking Latency in Advance
Building a model is exciting. Deploying it is urgent. But if latency requirements aren't clear before development, you might finish training a 2-second model that needs to serve in 100ms. Result: 6 months of optimization work. Solution: Define latency/throughput requirements before building the model.
Mistake 2: Training-Serving Skew
Model trained on preprocessed data, but serving code uses different preprocessing (different scaling, missing features, etc). Model makes poor predictions in production despite training accuracy of 95%. Solution: Use a feature store that both training and serving pipeline consume from.
Mistake 3: Not Testing Error Handling
Assuming models never crash. But what if input is malformed? Model out of memory? GPU fails? Production crashes silently or returns garbage. Solution: Input validation (Pydantic), exception handling, circuit breakers, fallback logic.
Mistake 4: Forgetting About Model Versioning
Deploy v2 of a model, traffic drops 30% (due to degradation), but can't rollback because no one documented v1. Solution: Always version models, keep previous versions, test v2 on 1% traffic before full rollout.
Mistake 5: Not Monitoring Data Drift
Model trained on 2023 data. In 2024, distribution shifts, accuracy drops to 60%. No one notices for 3 months. Solution: Monitor input distribution (via KL divergence), model predictions, actual labels (if available). Alert on drift.
Mistake 6: Deploying Bloated Models
Serving a 13B parameter model that could achieve 95% of the accuracy with 2B parameters (and be 6x faster). Solution: Early in development, test model size vs accuracy tradeoff. Distill if necessary.
Mistake 7: Ignoring Cold Start Time
Model takes 3 seconds to load into GPU memory. Every restart (deployments, crashes) has 3 second latency spike. With multiple pods restarting, cascades to 30 second downtime. Solution: Keep models warm, use persistent GPU memory, test startup time.
Mistake 8: Single Point of Failure
Only one instance of the model running. If it crashes, service is down until manual restart. Solution: Always run >= 3 replicas. Use auto-scaling. Have health checks.
Production Best Practices
Best practices are lessons learned by companies that got it wrong. Apply them from day one.
Development Phase
1. Define deployment requirements early: Latency (p50, p99), throughput (req/sec), accuracy, uptime SLA. Design model accordingly.
2. Build models for inference from start: Consider model size, quantizability, inference hardware. Avoid architectures that are hard to optimize.
3. Test with production-like inputs: Don't train on clean data, test on messy data. Simulate production traffic patterns.
4. Benchmark on target hardware: If deploying on T4 GPU, benchmark on T4 (not local GPU). Consider batch sizes, precision.
Deployment Phase
5. Use containers (Docker) for all deployments: Ensures reproducibility, simplifies dependency management, enables Kubernetes.
6. Version everything: Model artifacts, code, Docker images, configurations. Track lineage.
7. Implement health checks: Liveness (is the process alive?), readiness (can it handle traffic?), startup (how long to load model?)
8. Gradual rollout (canary deployment): Don't deploy v2 to 100% traffic at once. Deploy to 1%, then 5%, then 50%, with metrics comparison.
Production Phase
9. Monitor everything: Latency (p50, p99, p999), throughput, error rate, accuracy, resource utilization. Set up alerts.
10. Log for debugging: Log slow requests, errors, unusual inputs. Aggregrate with ELK or similar. But don't log raw predictions (privacy).
11. Implement circuit breakers: If model inference starts timing out, fail fast rather than cascading timeout.
12. Have a playbook for incidents: Model latency spikes, out of memory errors, accuracy degradation. Who to page? What to roll back?
Data & Governance
13. Monitor data drift: Track distribution of input features. Accuracy degradation often signals data shift.
14. Maintain audit trail: Log predictions for audit (legal requirement in some domains). Enable root cause analysis later.
15. Plan for retraining: Models degrade. Establish retraining cadence (weekly/monthly/quarterly) and automate it.
16. Document models: What data was it trained on? What performance guarantees? Known failure modes? Deprecation timeline?
Advanced Insights & State-of-the-Art
Beyond basics: cutting-edge techniques that leading companies use for competitive advantage.
Speculative Decoding for LLMs
Generating text token-by-token is slow. Speculative decoding uses a small fast model to generate draft tokens, then a large model verifies them in parallel. Result: Up to 2-3x faster text generation with no accuracy loss.
How it works: Small model generates 5 candidate tokens fast. Large model checks them all in parallel (much cheaper than generating them sequentially). If match, accept them all at once. If mismatch, use large model's choice. Net: 5 tokens generated in time of 1-2 large model calls.
Prefix Caching for Conversation
When chatting, you recompute attention for the same conversation history every turn. Prefix caching stores intermediate activations (KV cache) from previous turns. New query only computes attention on new tokens.
Performance gain: First turn of 1000-token conversation: 5 seconds. Second turn (new 50 tokens): 100ms (50x faster) instead of 5 seconds.
Model Merging & Ensemble
Train multiple models on different data or with different objectives, then merge them. A single merged model can match accuracy of ensemble while being faster.
Continuous Batching vs Micro-Batching
Continuous Batching (vLLM): Instead of waiting for a full batch before processing, start processing as soon as requests arrive. Each request generates tokens independently. When one finishes, replace with waiting request. GPU rarely sits idle.
Micro-batching: Process requests in tiny batches (size 1-4) rapidly. Slightly higher latency than continuous batching but simpler to implement.
Complete Code Examples
Real, runnable code for common deployment scenarios.
Example 1: End-to-End FastAPI + Docker
Example 2: ONNX Inference Optimization
Example 3: vLLM Serving
Hands-On Exercises
Apply what you've learned with practical exercises.
Exercise 1: Build Your First FastAPI Endpoint
Create a simple FastAPI app that loads a pretrained HuggingFace model and serves predictions via /predict endpoint. Test with curl.
Exercise 2: Containerize the API
Write a Dockerfile for your FastAPI app. Build the image, test it with `docker run`, and push to a registry.
Exercise 3: Quantize and Benchmark
Load a model in FP32, quantize to INT8 using ONNX, benchmark latency of both. Calculate speedup and accuracy change.
Exercise 4: Deploy to Kubernetes
Write a Kubernetes deployment manifest for your containerized model server. Apply it to a cluster and test with `kubectl port-forward`.
Interview Questions
Common questions asked in ML engineering interviews about deployment.
1. Profiling: Determine bottleneck. Is it model inference, I/O, preprocessing?
2. Model optimization: Quantization (FP32βINT8, 4x speedup). Distillation to smaller model. Pruning. Knowledge distillation.
3. Batching: Process requests in batches. 10x throughput (slight latency increase for last request).
4. Caching: Cache embeddings, feature preprocessing, model outputs if applicable.
5. Framework choice: vLLM for LLMs, Triton for complex pipelines. TensorFlow Serving for TF models.
6. Hardware: Use GPU. Faster GPU (V100 vs T4, 5x speedup). Multiple GPUs for ensemble.
7. Serving optimizations: Remove unnecessary preprocessing, minimize data copying, parallel model forward pass.
Typical result: 50-100x improvements combining several techniques.
Version Storage:
β’ Store model artifacts in model registry (MLflow, Hugging Face Hub) with version tags
β’ Include metadata: training date, accuracy, model hyperparameters, dependencies
Deployment Strategy:
β’ Never overwrite old model versions; always create new versions
β’ Test new version on staging environment first
β’ Canary deployment: Route 1% traffic to v2, monitor metrics, scale to 5%, 10%, 50%, 100%
β’ Keep previous versions readily available for rollback
Serving Strategy:
β’ Use model serving framework that supports multiple versions (Triton, TensorFlow Serving)
β’ Implement A/B testing: Split traffic between v1 and v2, compare accuracy/latency
β’ Automatic rollback if error rate increases >1% or latency >10%
Example: Deploy v1.5 to 1% traffic for 24 hours. If metrics are better, scale to 100%. If worse, automatic rollback to v1.4.
Latency Optimization (p99 < 100ms):
β’ Minimize time per single request
β’ Reduce model computation (quantization, pruning, smaller model)
β’ Reduce I/O (pre-load model, local inference)
β’ Trade throughput for latency (don't batch, or batch=1)
β’ Use fastest GPU, CPU processor
Throughput Optimization (max req/sec):
β’ Maximize requests per second
β’ Batch requests aggressively (batch_size=128), slight latency increase
β’ Parallel processing (multi-GPU, multi-model)
β’ Can tolerate higher latency per request if total throughput increases
Example: Recommendation system needs latency. Spam detection can batch (can tolerate 100ms latency for throughput).
Serving Metrics (Real-time):
β’ Latency: p50, p99, p999 (not just mean)
β’ Throughput: requests per second
β’ Error rate: % of requests failing
β’ Resource utilization: GPU memory, CPU, bandwidth
Model Metrics (Batch):
β’ Accuracy: True positive rate, false positive rate on holdout test set
β’ Data drift: Distribution of input features vs training distribution (KL divergence)
β’ Prediction distribution: Are outputs changing unexpectedly?
Infrastructure Metrics:
β’ Pod health: % replicas up, deployment success rate
β’ Kubernetes health: Node availability, persistent storage
Tools:
β’ Prometheus for metrics collection
β’ Grafana for dashboards
β’ ELK Stack for logs
β’ Weights & Biases or similar for model-specific monitoring
Alerting:
β’ Alert if p99 latency > 200ms
β’ Alert if error rate > 1%
β’ Alert if data drift detected
β’ Alert if accuracy drops > 2%
Problem: First request to a container incurs 1-5 second latency penalty for loading model into GPU memory. In auto-scaled systems with frequent deployment, cold starts are common and visible to users.
Solutions:
1. Warm Up on Startup:
β’ Send dummy requests in container startup script
β’ Load model and run inference before marking as ready
β’ Kubernetes readiness probe won't pass until warm
2. Model Preloading:
β’ Use Init Container in Kubernetes to load model from shared storage
β’ Pin model weights in GPU memory (don't evict)
3. Prevent Scale Down:
β’ Minimum replicas > 0 (keep at least 1 warm instance always running)
β’ Don't scale to zero, even during low traffic
4. Fast Startup:
β’ Use smaller models for initialization (distilled version)
β’ Cache compiled CUDA kernels
Tradeoff: Idle cost vs cold start latency. Usually worth keeping warm replica running.
Batch Inference:
β’ Process large dataset (1M+ samples) at once
β’ Latency not critical (can wait hours)
β’ Throughput critical (maximize GPU utilization)
β’ Example: Daily recommendation computation, overnight fraud scoring
β’ Implementation: Spark, Kubernetes batch jobs, SageMaker Batch Transform
Online Serving:
β’ Process single request in response to API call
β’ Latency critical (must respond in 100ms)
β’ Throughput secondary (though important at scale)
β’ Example: Real-time search ranking, recommendation ranking
β’ Implementation: FastAPI, Triton, vLLM
Mixed Approach (Lambdas Architecture):
β’ Use batch inference for complex, heavy models overnight
β’ Cache results (e.g., Redis)
β’ Online serving looks up cached results
β’ Combine batch + online for best of both worlds
The Problem (Training-Serving Skew):
β’ Training pipeline: Complex feature engineering in Spark
β’ Serving pipeline: Different code (maybe different language), slightly different logic
β’ Result: Model sees different features at train time vs serve time
β’ Impact: Accuracy drops from 95% β 80% in production
Solution 1: Feature Store
β’ Centralized system for feature definitions and computation
β’ Both training and serving pipelines consume from feature store
β’ Ensures identical features
β’ Examples: Tecton, Feast, Vertex AI Feature Store
Solution 2: Unified Code
β’ Write feature engineering in language-agnostic format (e.g., SQL)
β’ Compile to both training and serving
β’ Use same preprocessing for training and serving
Solution 3: Versioned Features
β’ Version feature engineering code
β’ When training model, record feature version used
β’ At serving time, use same feature version
Solution 4: Testing
β’ Compare features computed in training vs serving on same data
β’ Alert if differences > threshold
Requirements Analysis:
β’ 100 models, different sizes (1MB - 50GB)
β’ Different latency needs (some need <50ms, some tolerate 1s)
β’ Shared GPU resources
β’ Cost optimization needed
Architecture:
1. Model Registry:
β’ MLflow: Store all 100 models with metadata (size, latency profile, accuracy)
β’ Tag each model with SLA requirements
2. Categorization by Latency:
β’ <50ms: Fast models (25 models, 2GB each) β Triton on expensive GPUs (V100)
β’ 50-200ms: Medium models (50 models) β vLLM or TF Serving on T4/A100
β’ >200ms: Slow models (25 models) β Batch inference, Kubernetes jobs
3. Serving Strategy:
β’ API Gateway: Routes requests to appropriate serving infrastructure
β’ Model Router: Determines which serving engine for each model
β’ Load Balancer: Distributes to instances
4. GPU Sharing:
β’ Triton Ensemble: Run multiple small models on 1 GPU
β’ Memory isolation: Limit GPU memory per model
β’ Context switching: Switch between models based on traffic
5. Cost Optimization:
β’ Spot instances for batch inference (25% cost)
β’ Quantization for all models (reduce GPU requirements)
β’ Idle model unloading (free up GPU after N minutes)
6. Monitoring:
β’ Per-model metrics dashboard
β’ Alert if any model violates SLA
β’ Track GPU utilization per model
Frequently Asked Questions
Summary & Key Takeaways
Model deployment is both art and science. Here are the essential principles:
Core Principles
1. Latency Matters: 100ms might not matter for batch inference, but it's critical for real-time systems. Measure it, optimize it, monitor it.
2. Reliability > Accuracy: A model that crashes is worthless. Invest in monitoring, alerts, and failover mechanisms before optimizing model accuracy.
3. Cost Discipline: GPU costs scale with traffic. Quantization, batching, and optimization save millions. Build cost optimization into model development.
4. Automation: Manual deployments are error-prone and slow. Invest in CI/CD pipelines. Models should deploy with same rigor as code.
5. Observability: You can't optimize what you don't measure. Instrument everything: latency, accuracy, resource usage, errors.
The Deployment Checklist
Before deploying to production, verify:
β Latency/throughput requirements defined and met
β Model quantized (INT8 or FP16) for production
β Containerized (Docker) with health checks
β Tested locally and in staging
β Input validation and error handling implemented
β Monitoring and alerting configured
β Model versioning strategy in place
β Rollback procedure documented
β Canary deployment plan created
β On-call runbook written
Next Steps
You now understand model deployment. To master it:
- Build: Deploy a simple model using FastAPI + Docker. Get comfortable with basics.
- Optimize: Quantize it, benchmark latency improvements, measure cost savings.
- Scale: Deploy to Kubernetes, implement auto-scaling, monitor in production.
- Advanced: Explore vLLM, Triton, multi-model serving, feature stores.
- Production: Lead a real model deployment at your company. Learn from mistakes.
Remember: The best model is the one serving users in production. Deployment isn't the final step β it's the beginning of the ML system's life. Invest accordingly.
Resources & Further Learning
Official Frameworks & Documentation
- FastAPI:
https://fastapi.tiangolo.comβ Modern Python web framework - vLLM:
https://docs.vllm.aiβ LLM serving with continuous batching - TensorFlow Serving:
https://www.tensorflow.org/tfx/guide/servingβ Production TF model serving - Triton Inference Server:
https://docs.nvidia.com/tritonβ Multi-model multi-framework serving - Kubernetes:
https://kubernetes.io/docsβ Container orchestration - Docker:
https://docs.docker.comβ Containerization platform - MLflow:
https://mlflow.orgβ Model registry and experiment tracking
Recommended Books & Papers
- "Designing Machine Learning Systems" by Chip Huyen β Comprehensive guide to ML systems including deployment
- "Production-Ready Microservices" by Susan Fowler β How to build reliable production systems
- "Efficient Inference for Large Language Models" β Recent papers on serving LLMs efficiently
- vLLM Paper: "vLLM: Easy, Fast, and Cheap LLM Serving" β Explains continuous batching
Courses & Tutorials
- Full Stack LLM Bootcamp (UC Berkeley) β Covers deployment of LLMs
- Stanford CS329S β ML systems design course
- Hugging Face Course:
huggingface.co/courseβ NLP and model deployment - DeepLearning.AI: Short courses on MLOps, LLMOps
Tools & Platforms
- Model Registries: MLflow, Hugging Face Hub, NVIDIA NGC
- Serving Frameworks: FastAPI, vLLM, Triton, TensorFlow Serving, KServe
- Orchestration: Kubernetes, Docker Swarm, Docker Compose
- Monitoring: Prometheus, Grafana, DataDog, New Relic
- Feature Stores: Tecton, Feast, Databricks Feature Store
- Optimization: ONNX, TensorRT, PyTorch's torch.compile
Community & Blogs
- Chip Huyen's Blog:
huyenchip.comβ ML systems and deployment insights - vLLM Team Blogs: Latest LLM serving techniques
- NVIDIA Developer Blog: GPU optimization and TensorRT
- Kubernetes Community:
kubernetes.io/community - Fast.ai Forums: Active community discussion on deployment