Talent.com
qpisemi
AI20P Library Engineer, Machine Learning Accelerationqpisemi • Bengaluru / Bangalore, India
AI20P Library Engineer, Machine Learning Acceleration

AI20P Library Engineer, Machine Learning Acceleration

qpisemi • Bengaluru / Bangalore, India
25 days ago
Job description

AI20P Library Engineer, Machine Learning Acceleration

About the Role

We are looking for an experienced AI20P Library Engineer to bridge the gap between cutting-edge AI research and our hardware accelerators. In this role, you will design, develop, and optimize high-performance software libraries, kernels, and parallel compute runtimes for distributed AI20P environments. You will empower ML researchers and cloud developers to train and deploy frontier AI models at unprecedented scales.

Key Responsibilities

Kernel Development & Optimization: Design, implement, and tune high- performance custom operators and mathematical kernels specifically for AI20P architectures (assembly or intrinsic level).

Distributed Computing Runtimes: Build software abstractions and libraries that manage multi-host setups, sharding, and multi-dimensional parallelization (e.g., Megatron-LM style tensor/pipeline parallelism).

Communication Primitives: Design and optimize custom collective communication algorithms (AllReduce, AllGather) to minimize latency and maximize throughput over high-speed interconnects.

Performance Tuning: Profile distributed training workloads to identify and eliminate memory bottlenecks, network stalls, and suboptimal hardware utilization.

Cross-Functional Collaboration: Partner with compiler developers (XLA), hardware architects, and research scientists to co-design the future software- hardware ecosystem.

Qualifications

Education: B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a highly quantitative field.

Programming: Exceptional proficiency in C++ and Python.

ML Ecosystem: Hands-on experience with at least one major AI/ML framework (e.g., JAX, PyTorch).

Systems Knowledge: Strong understanding of concurrent computations, memory hierarchies (HBM, cache), and hardware accelerators.

Communication: Excellent cross-functional communication skills to translate complex research needs into production-ready software architecture.

Preferred Qualifications (Nice-to-Have)

Experience working with compiler construction or optimizing code generation for hardware.

Familiarity with cloud-based cluster managers (e.g., Kubernetes, SLURM) for executing large-scale distributed ML workloads.

Prior contributions to optimizing Large Language Models (LLMs) or multimodal architectures for low-precision inference and training.


Skills Required
Performance Tuning, Kernel Development, Python
Create a job alert for this search

AI20P Library Engineer, Machine Learning Acceleration • Bengaluru / Bangalore, India