Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Production

  • Packaging models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Managing versioning and storage requirements

Serving Models on Kubernetes

  • An overview of inference server architectures
  • Deploying TensorFlow Serving and TorchServe
  • Establishing efficient model endpoints

Optimizing Inference Performance

  • Implementing effective batching strategies
  • Handling concurrent requests efficiently
  • Tuning for optimal latency and throughput

Autoscaling ML Workloads

  • Utilizing the Horizontal Pod Autoscaler (HPA)
  • Leveraging the Vertical Pod Autoscaler (VPA)
  • Applying Kubernetes Event-Driven Autoscaling (KEDA)

GPU Provisioning and Resource Management

  • Configuring dedicated GPU nodes
  • An overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Model Rollout and Release Management

  • Implementing blue/green deployment patterns
  • Using canary rollout strategies
  • Conducting A/B testing for model evaluation

Monitoring and Observability in Production ML

  • Tracking key metrics for inference workloads
  • Best practices for logging and tracing
  • Creating dashboards and setting up alerting

Security and Reliability Best Practices

  • Securing model endpoints against threats
  • Implementing network policies and access controls
  • Ensuring high availability of services

Summary and Future Directions

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Familiarity with core Kubernetes concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Testimonials (4)

Upcoming Courses

Related Categories