Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Performance: Concepts and Metrics
- Analysis of latency, throughput, power consumption, and resource utilization
- Distinguishing between system-level and model-level bottlenecks
- Profiling strategies tailored for inference versus training workflows
Profiling Strategies on Huawei Ascend
- Leveraging CANN Profiler and MindInsight for deep insights
- Diagnostic techniques for kernels and operators
- Understanding offload patterns and memory mapping mechanisms
Profiling Strategies on Biren GPU
- Utilizing Biren SDK features for performance monitoring
- Optimizing kernel fusion, memory alignment, and execution queues
- Implementing power and temperature-aware profiling
Profiling Strategies on Cambricon MLU
- Employing BANGPy and Neuware performance utilities
- Gaining kernel-level visibility and interpreting diagnostic logs
- Integrating MLU profilers with various deployment frameworks
Advanced Graph and Model Optimization
- Strategies for graph pruning and quantization
- Techniques for operator fusion and restructuring computational graphs
- Standardizing input sizes and fine-tuning batch processing
Memory and Kernel Optimization Techniques
- Enhancing memory layout and reuse efficiency
- Implementing efficient buffer management across different chipsets
- Applying platform-specific kernel tuning methods
Best Practices for Cross-Platform Performance
- Achieving performance portability through abstraction strategies
- Developing shared tuning pipelines suitable for multi-chip environments
- Case study: optimizing an object detection model across Ascend, Biren, and MLU
Conclusion and Future Directions
Requirements
- Prior experience managing AI model training or deployment pipelines
- Solid grasp of GPU/MLU computing principles and model optimization techniques
- Foundational knowledge of performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours