AI/ML Engineer - LLM & Enterprise AI Runtime
Location: New Delhi, India (On-site)
Employment Type: Full-time
About QubeLabs
QubeLabs is building the next generation of Enterprise AI Systems that transform workforce operations and intelligence for the financial services industry. Our platform combines conversational AI, agentic workflow automation, proprietary language models and enterprise intelligence to transform how financial services operate.
Our product "QubeLabs Workmate" is purpose-built for banks, NBFCs, MFIs, wealth management firms, insurance companies and fintechs across India and Europe to improve enterprise efficiency, productivity and customer experiences at scale.
We're looking for an exceptional AI/ML Engineer who is passionate about building production-grade Enterprise AI systems and enjoys working at the intersection of Large Language Models, real-time AI runtimes, enterprise intelligence and AI orchestration.
Role Overview
As an AI/ML Engineer, you will build and optimize the Enterprise AI Runtime powering QubeLabs Workmate. You will work closely with founders, product teams, Full Stack Engineers and Speech AI Engineers to develop real-time conversational AI, enterprise copilots and intelligent workflow automation.
This role is ideal for engineers with deep expertise in Large Language Models, AI orchestration, retrieval systems and low-latency inference. You will leverage modern AI engineering tools such as Cursor, Claude Code, GitHub Copilot and similar AI development assistants to accelerate research and development while maintaining production-grade AI systems. You will contribute to the core intelligence layer powering conversational AI, reasoning, tool calling, enterprise knowledge retrieval and agentic workflows.
Key Responsibilities
Enterprise AI Runtime
- Design and develop the Enterprise AI Runtime for conversational AI, enterprise copilots and intelligent workflow automation.
- Build AI orchestration capabilities including context management, memory, tool calling and model routing.
- Develop low-latency inference pipelines for real-time voice conversations and AI copilots.
- Design intelligent execution strategies for enterprise workflows and AI services.
- Optimize AI runtime for scalability, observability, reliability and production deployment.
LLM Engineering
- Fine-tune and optimize open-source Large Language Models and multimodal models using Supervised Fine-Tuning (SFT), instruction tuning and preference optimization techniques.
- Build enterprise reasoning, structured output generation and domain-specific AI capabilities.
- Develop Retrieval-Augmented Generation (RAG), semantic search and enterprise knowledge retrieval systems.
- Improve model accuracy through prompt engineering, grounding, hallucination reduction and evaluation frameworks.
- Develop multilingual reasoning and enterprise domain adaptation capabilities.
AI Platform Development
- Develop model serving APIs, reusable inference services and AI SDKs.
- Optimize model serving using quantization, batching, caching and GPU acceleration.
- Build AI evaluation pipelines, benchmarking frameworks and performance monitoring.
- Collaborate with Speech AI Engineers to integrate ASR, NLU, NER and TTS pipelines with LLM-powered conversational systems.
Engineering Excellence
- Build reproducible AI pipelines and MLOps workflows.
- Participate in AI architecture, model evaluation and system design discussions.
- Leverage AI-assisted engineering tools to improve research and development productivity.
- Follow engineering best practices for experimentation, versioning, deployment, monitoring and continuous or Master's degree in Computer Science, Artificial Intelligence, Machine Learning or a related discipline.
- Graduates from Tier-1 institutions are years of experience in AI/ML Engineering, Applied AI or Machine Learning.
- Experience building production-grade AI systems is preferred.
- Experience with Large Language Models, conversational AI, enterprise AI platforms or real-time AI systems is highly desirable.
- Experience in model fine-tuning, inference optimization, AI orchestration or Retrieval-Augmented Generation (RAG) will be an added advantage.
- Candidates should demonstrate strong proficiency in AI-assisted engineering and the ability to significantly improve research and development productivity using modern AI engineering tools while maintaining production-grade AI systems.
Required SkillsLarge Language Models :Open-source LLMs (Llama, Qwen, Gemma, Mistral or equivalent), Multimodal Models, Hugging Face Transformers, vLLM, SGLang, TensorRT-LLM, Prompt Engineering, Structured Output Generation, Function Calling and Tool Calling.
Enterprise AI Runtime:AI Orchestration, Multi-Agent Systems, Context Management, Session Memory, Conversation Runtime, Retrieval-Augmented Generation (RAG), Hybrid Search, Vector Databases (Qdrant, Pinecone or equivalent), Model Context Protocol (MCP), Agent Frameworks (LangGraph, AutoGen, CrewAI or equivalent).
Model Engineering: Supervised Fine-Tuning (SFT), Instruction Tuning, Preference Optimization (DPO or equivalent), Model Evaluation, Benchmarking, Hallucination Reduction, Grounding, Quantization, GPU Optimization and Inference Python, PyTorch, FastAPI, Hugging Face, REST APIs, Docker, Redis, PostgreSQL, Git and Linux.
Cloud & MLOps: Docker, Kubernetes, AWS/Azure/GCP, MLflow, Weights & Biases, CI/CD, Model Deployment, Monitoring and Development
Working knowledge of:
- OpenAI, Anthropic, Google Gemini and Open Source AI Models
- AI SDKs and Inference Frameworks
- AI Evaluation Frameworks
- Streaming APIs
- Enterprise AI Architectures
- Working knowledge of Speech AI pipelines (ASR, NLU, NER and TTS) and experience integrating speech AI with LLM-powered conversational systems will be an added advantage.
- Mandatory: Hands-on proficiency using Cursor, Claude Code, GitHub Copilot or equivalent AI engineering assistants as part of day-to-day AI Attributes
- Strong analytical and problem-solving skills.
- Passion for AI research and applied engineering.
- Curious about emerging AI technologies.
- Strong software engineering mindset.
- Detail-oriented with excellent execution capability.
- Self-driven with a high sense of ownership.
- Excellent communication and collaboration skills.
- Ability to thrive in a fast-paced startup environment.
What Success Looks Like
Within the first year, you will:
- Build the Enterprise AI Runtime powering QubeLabs Workmate.
- Deliver production-ready LLM capabilities for conversational AI and enterprise copilots.
- Develop scalable AI orchestration, context management and retrieval systems.
- Optimize inference pipelines for low-latency, high-performance enterprise deployments.
- Contribute significantly to the architecture and intelligence foundation of QubeLabs' next-generation Enterprise AI Platform.
Why Join QubeLabs?
- Build the future of Enterprise AI Systems.
- Work directly with founders and leadership.
- Develop next-generation AI systems for leading financial institutions.
- Work on cutting-edge technologies including Conversational AI, Agentic AI, Enterprise Intelligence, Real-Time AI Copilots and Sovereign AI.
- Exposure to global markets across India and Europe.
- High ownership, rapid learning and accelerated career growth.
Preferred Candidate Profile
- Graduate from Tier-1 engineering institutions (IITs, IIITs, IISc, IIIT Hyderabad, BITS Pilani, NITs, NSUT, DTU or equivalent).
- Strong GitHub, Hugging Face or open-source portfolio demonstrating production AI work.
- Experience building real-time conversational AI, enterprise copilots or AI-native platforms.
- Experience with low-latency inference, distributed model serving, AI orchestration, GPU optimization or multimodal AI systems is highly preferred.
- Contributions to open-source AI projects, research publications or enterprise AI products will be an added advantage.
- Passion for AI, technology and building next-generation Enterprise AI Systems.
QubeLabs is an equal opportunity employer committed to building a diverse and inclusive workplace where innovation thrives (ref:hirist.tech)