Job Description:
- Research, design, and train AI models for real-time performance evaluation across domains (music first, then dance/speech/chess).
- Implement and optimize deep learning architectures for audio and/or visual understanding (CNNs, RNNs, Transformers).
- Work closely with Audio/Vision Engineers to build data pipelines for clean, real-time feature extraction (spectrograms, keypoints, pose sequences).
- Collaborate with SMEs to define performance quality metrics and label datasets.
- Develop evaluation frameworks to quantify model accuracy vs. expert feedback.
- Experiment with cross-modal fusion (audio + vision) for synchronized analysis in future domains like dance.
- Optimize models for low-latency inference on web/mobile devices (ONNX, TensorRT, TF Lite).
- Document research findings, prototype outcomes, and contribute to internal knowledge-sharing.
Required Skills & Experience:
- 3+ years of hands-on experience in Machine Learning / Deep Learning (PyTorch, TensorFlow).
- Strong mathematical foundation in signal processing, time-series analysis, and statistics.
- Proven experience with audio or visual data music, speech, motion, or similar perceptual domains.
- Familiarity with MIR (Music Information Retrieval) or Computer Vision tasks:
1. Pitch detection, beat tracking, timbre classification, speech analysis.
2. Pose estimation, gesture recognition, or motion tracking.
- Experience with model optimization and deployment (TorchScript, ONNX, TensorRT).
- Strong Python skills and familiarity with libraries such as NumPy, pandas, Librosa, Essentia, OpenCV, or MediaPipe.
Preferred Qualifications:
- Candidates who have completed Bachelor's or Master's from IIT.
(ref:hirist.tech)