Get in Touch

Course Outline

Overview of Biren GPU Architecture

  • Biren platform overview and key use cases
  • Hardware components: cores, memory structure, and compute clusters
  • Architectural comparison with NVIDIA and AMD GPUs

Establishing the Biren Programming Environment

  • Installing the Biren SDK and runtime components
  • Exploring the toolchain and compiler architecture
  • Structuring basic projects and managing the build process

Developing Applications with the Biren Stack

  • Thread and block organizational models
  • Managing memory and handling data transfers
  • Designing kernels and implementing launch patterns

Migrating from CUDA to Biren

  • Strategies for translating CUDA code
  • Mapping common APIs and adapting functionalities
  • Hands-on code conversion labs and practical exercises

Debugging and Profiling Strategies

  • Leveraging Biren’s debugger and profiler tools
  • Identifying performance bottlenecks
  • Analyzing memory access patterns for optimization

Advanced Optimization Techniques

  • Managing thread scheduling and instruction pipelining
  • Utilizing loop unrolling and shared memory effectively
  • Fine-tuning kernels for maximum throughput

Real-World Case Studies and Applications

  • Training models using Biren accelerators
  • Porting and profiling vision or NLP models
  • Benchmarking performance against CUDA/NVIDIA ecosystems

Summary and Recommended Next Steps

Requirements

  • Solid understanding of GPU architecture and parallel processing concepts
  • Proficiency in GPU programming environments such as CUDA, OpenCL, or similar frameworks
  • Working knowledge of deep learning frameworks like PyTorch or TensorFlow

Target Audience

  • HPC Developers
  • AI Infrastructure Engineers
  • Performance Optimization Specialists
 21 Hours

Testimonials (2)

Upcoming Courses

Related Categories