

Search by job, company or skills

Job Title: MLOps Lead - Enterprise AI Platform
Location: Bangalore (Hybrid)
Responsibilities:
Strategic Platform Architecture:
• Lead the architectural vision, design, and continuous evolution of the Platform, ensuring alignment with business objectives, security standards, and scalability requirements.
• Drive the adoption and integration of open-source MLOps tools (Kubeflow, MLflow, Feast, KServe, Alibi-Detect, Evidently AI, Spark, etc.) into a cohesive, production-ready enterprise solution. • Define platform standards, best practices, and architectural patterns for MLOps development and operations.
Technical Leadership & Implementation Oversight:
• Act as the primary technical authority and lead for the MLOps initiative, guiding both DevOps/Platform and MLOps/Data Science teams through the phased development plan.
• Oversee the implementation of core platform components, ensuring robust integration, performance, and adherence to architectural blueprints.
• Provide expert guidance on Kubernetes-native MLOps practices, distributed computing for ML (Spark, Kubeflow Training Operators), and model serving strategies (KServe).
Enterprise Security, Governance & Multi-Tenancy:
Architect and oversee the implementation of enterprise-grade security features including SSO (Keycloak), secrets management (HashiCorp Vault), and fine-grained access control (Kubernetes RBAC, OPA Gatekeeper) for data and platform resources.
• Design and enforce multi-tenancy models that provide strong isolation, resource governance, and secure data access for internal teams and external customers.
• Ensure the platform meets stringent compliance requirements through comprehensive audit logging, tracing (Fluentd, ELK/OpenSearch, Prometheus/Grafana), and data lineage considerations.
ML Lifecycle & Data Management Expertise:
• Architect and integrate a robust Feature Store (tool like Feast) for consistent feature engineering, management, and serving across training and inference.
• Lead the integration of MLflow for experiment tracking, model versioning, and a centralized model registry.
• Design and implement comprehensive model monitoring solutions, including data drift and model quality detection (Alibi-Detect/Evidently AI), with integrated alerting.
Developer Experience & Customization:
• Champion the developer experience for data scientists, ensuring ease of use, self-service capabilities, and efficient workflows (e.g., automated namespace provisioning, notebook environment management).
• Provide architectural guidance for building a custom, branded UI layer on top of the open source components, enhancing usability and aligning with product offerings.
Collaboration & Mentorship:
• Collaborate extensively with Data Science, DevOps, Security, Product Management, and Business stakeholders to gather requirements, communicate technical vision, and drive platform adoption.
• Mentor and upskill engineering teams in MLOps best practices, cloud-native development, and advanced ML techniques.
Required Skills & Expertise:
• 7+ years of progressive experience in software engineering, data engineering, or MLOps, with at least 5 years in a lead or architect role focused on building and managing production of large-scale ML platforms.
• Expert-level proficiency with Kubernetes and its ecosystem (operators, CRDs, Helm, networking, storage).
• Experience in building/managing ML platform tools such as MLflow , Kubeflow, Airflow, SageMaker, Vertex AI, or Azure Machine Learning.
• Deep hands-on experience with Kubeflow (Pipelines, Notebooks, Training Operators, KServe) in production environments.
• Extensive experience with MLflow for experiment tracking, model registry, and model lifecycle management.
Proven expertise in designing and implementing Feature Stores (e.g., Feast) for both online and offline serving.
• Strong background in distributed data processing technologies like Apache Spark/PySpark, especially on Kubernetes.
• Architectural experience with enterprise security solutions including SSO (Keycloak, OAuth/OIDC), secrets management (HashiCorp Vault), and policy enforcement (Kubernetes RBAC, OPA Gatekeeper).
• Demonstrated ability to implement comprehensive monitoring and observability stacks (Prometheus, Grafana, ELK/OpenSearch, Fluentd, Jaeger) for platform health and ML model performance/drift (Alibi-Detect, Evidently AI).
• Proficiency in Python and experience with major ML/Deep Learning frameworks (TensorFlow, PyTorch, Scikit-learn).
• Experience with cloud-native storage solutions (e.g., MinIO, S3, GCS) and open table formats (Iceberg, Delta Lake).
• Excellent communication, leadership, and interpersonal skills with the ability to influence technical direction and drive complex initiatives across multiple teams.
RAKUTEN SHUGI PRINCIPLES: Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
• Always improve, always advance. Only be satisfied with complete success - Kaizen.
• Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
• Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
• Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
• Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.
Job ID: 151586549
Skills:
Tensorflow, Git, Pytorch, MLops, Docker, Flask, FastAPI, Python, Kubernetes, AWS, Airflow, vector databases, embedding models, Sagemaker, Streamlit