Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
CANN Optimisation Capabilities Overview
- How inference performance is managed within CANN
- Optimisation targets for edge and embedded AI systems
- Comprehending AI Core utilisation and memory allocation
Utilising the Graph Engine for Analysis
- Introduction to the Graph Engine and execution pipeline
- Visualising operator graphs and runtime metrics
- Refining computational graphs for optimal performance
Profiling Tools and Performance Metrics
- Applying the CANN Profiling Tool (profiler) for workload analysis
- Evaluating kernel execution time and identifying bottlenecks
- Memory access profiling and tiling strategies
Custom Operator Development with TIK
- Overview of TIK and its operator programming model
- Implementing custom operators using the TIK DSL
- Testing and benchmarking operator performance
Advanced Operator Optimisation with TVM
- Introduction to TVM integration with CANN
- Auto-tuning strategies for computational graphs
- Determining when and how to switch between TVM and TIK
Memory Optimisation Techniques
- Managing memory layout and buffer placement
- Techniques to minimise on-chip memory consumption
- Best practices for asynchronous execution and data reuse
Real-World Deployment and Case Studies
- Case study: Performance tuning for smart city camera pipelines
- Case study: Optimising the inference stack for autonomous vehicles
- Guidelines for iterative profiling and continuous improvement
Summary and Next Steps
Requirements
- A solid grasp of deep learning model architectures and training workflows
- Practical experience with model deployment using CANN, TensorFlow, or PyTorch
- Proficiency in Linux CLI, shell scripting, and Python programming
Target Audience
- AI performance engineers
- Specialists in inference optimisation
- Developers working on edge AI or real-time systems
14 Hours