Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Custom Operator Development
- Rationale for custom operators: exploring use cases and architectural constraints
- Examining CANN runtime architecture and key operator integration touchpoints
- Contextualizing TBE, TIK, and TVM within the broader Huawei AI ecosystem
Implementing Low-Level Operators with TIK
- Grasping the TIK programming model and its supported API surface
- Managing memory and applying tiling strategies effectively in TIK
- The complete workflow: creating, compiling, and registering custom ops within CANN
Validation and Testing of Custom Operators
- Conducting unit and integration tests for operators within the execution graph
- Identifying and resolving kernel-level performance bottlenecks
- Visualizing operator execution flows and buffer memory behavior
Scheduling and Optimization via TVM
- Overview of TVM as a compiler framework for tensor operations
- Authoring schedules for custom operators using TVM
- Applying TVM for tuning, benchmarking, and code generation targeting Ascend
Framework and Model Integration
- Registering custom operators for compatibility with MindSpore and ONNX
- Ensuring model integrity and verifying fallback mechanisms
- Supporting multi-operator graphs that utilize mixed precision
Case Studies and Advanced Optimization Techniques
- Case study: Achieving high-efficiency convolution for small input shapes
- Case study: Optimizing attention operators with memory-awareness
- Best practices for deploying custom operators across diverse devices
Recap and Future Pathways
Requirements
- In-depth understanding of AI model internals and operator-level computational logic
- Proficiency in Python and Linux-based development environments
- Working knowledge of neural network compilers or graph-level optimization tools
Target Audience
- Compiler engineers engaged in AI toolchain development
- Systems developers specializing in low-level AI performance optimization
- Engineers developing custom operators or addressing novel AI workloads
14 Hours