Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Production Deployment

  • Encapsulating models using Docker
  • Exporting models from TensorFlow and PyTorch ecosystems
  • Best practices for versioning and storage management

Serving Models within Kubernetes

  • An introduction to inference server architectures
  • Deploying TensorFlow Serving and TorchServe instances
  • Configuration and management of model endpoints

Optimizing Inference Performance

  • Implementing effective batching strategies
  • Handling concurrent requests efficiently
  • Tuning for optimal latency and throughput

Autoscaling Machine Learning Workloads

  • Utilizing the Horizontal Pod Autoscaler (HPA)
  • Leveraging the Vertical Pod Autoscaler (VPA)
  • Implementing Kubernetes Event-Driven Autoscaling (KEDA)

Managing GPUs and Resource Allocation

  • Configuration of GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Strategies for Model Rollout and Release

  • Implementing blue/green deployment patterns
  • Executing canary rollout workflows
  • Conducting A/B testing for model evaluation

Monitoring and Observability for Production ML

  • Tracking key metrics for inference workloads
  • Establishing robust logging and tracing practices
  • Configuring dashboards and alerting mechanisms

Security and Reliability Best Practices

  • Protecting model endpoints from unauthorized access
  • Enforcing network policies and access controls
  • Safeguarding high availability standards

Course Summary and Future Directions

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience working with Python-based machine learning models
  • Foundational knowledge of Kubernetes principles

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories