Search by job, company or skills

  • Posted 19 days ago
  • Be among the first 10 applicants

Job Description

Job Title: MLOps Lead - Enterprise AI Platform

Location: Bangalore (Hybrid)

Responsibilities:

Strategic Platform Architecture:

• Lead the architectural vision, design, and continuous evolution of the Platform, ensuring alignment with business objectives, security standards, and scalability requirements.

• Drive the adoption and integration of open-source MLOps tools (Kubeflow, MLflow, Feast, KServe, Alibi-Detect, Evidently AI, Spark, etc.) into a cohesive, production-ready enterprise solution. • Define platform standards, best practices, and architectural patterns for MLOps development and operations.

Technical Leadership & Implementation Oversight:

• Act as the primary technical authority and lead for the MLOps initiative, guiding both DevOps/Platform and MLOps/Data Science teams through the phased development plan.

• Oversee the implementation of core platform components, ensuring robust integration, performance, and adherence to architectural blueprints.

• Provide expert guidance on Kubernetes-native MLOps practices, distributed computing for ML (Spark, Kubeflow Training Operators), and model serving strategies (KServe).

Enterprise Security, Governance & Multi-Tenancy:

Architect and oversee the implementation of enterprise-grade security features including SSO (Keycloak), secrets management (HashiCorp Vault), and fine-grained access control (Kubernetes RBAC, OPA Gatekeeper) for data and platform resources.

• Design and enforce multi-tenancy models that provide strong isolation, resource governance, and secure data access for internal teams and external customers.

• Ensure the platform meets stringent compliance requirements through comprehensive audit logging, tracing (Fluentd, ELK/OpenSearch, Prometheus/Grafana), and data lineage considerations.

ML Lifecycle & Data Management Expertise:

• Architect and integrate a robust Feature Store (tool like Feast) for consistent feature engineering, management, and serving across training and inference.

• Lead the integration of MLflow for experiment tracking, model versioning, and a centralized model registry.

• Design and implement comprehensive model monitoring solutions, including data drift and model quality detection (Alibi-Detect/Evidently AI), with integrated alerting.

Developer Experience & Customization:

• Champion the developer experience for data scientists, ensuring ease of use, self-service capabilities, and efficient workflows (e.g., automated namespace provisioning, notebook environment management).

• Provide architectural guidance for building a custom, branded UI layer on top of the open source components, enhancing usability and aligning with product offerings.

Collaboration & Mentorship:

• Collaborate extensively with Data Science, DevOps, Security, Product Management, and Business stakeholders to gather requirements, communicate technical vision, and drive platform adoption.

• Mentor and upskill engineering teams in MLOps best practices, cloud-native development, and advanced ML techniques.

Required Skills & Expertise:

• 7+ years of progressive experience in software engineering, data engineering, or MLOps, with at least 5 years in a lead or architect role focused on building and managing production of large-scale ML platforms.

• Expert-level proficiency with Kubernetes and its ecosystem (operators, CRDs, Helm, networking, storage).

• Experience in building/managing ML platform tools such as MLflow , Kubeflow, Airflow, SageMaker, Vertex AI, or Azure Machine Learning.

• Deep hands-on experience with Kubeflow (Pipelines, Notebooks, Training Operators, KServe) in production environments.

• Extensive experience with MLflow for experiment tracking, model registry, and model lifecycle management.

Proven expertise in designing and implementing Feature Stores (e.g., Feast) for both online and offline serving.

• Strong background in distributed data processing technologies like Apache Spark/PySpark, especially on Kubernetes.

• Architectural experience with enterprise security solutions including SSO (Keycloak, OAuth/OIDC), secrets management (HashiCorp Vault), and policy enforcement (Kubernetes RBAC, OPA Gatekeeper).

• Demonstrated ability to implement comprehensive monitoring and observability stacks (Prometheus, Grafana, ELK/OpenSearch, Fluentd, Jaeger) for platform health and ML model performance/drift (Alibi-Detect, Evidently AI).

• Proficiency in Python and experience with major ML/Deep Learning frameworks (TensorFlow, PyTorch, Scikit-learn).

• Experience with cloud-native storage solutions (e.g., MinIO, S3, GCS) and open table formats (Iceberg, Delta Lake).

• Excellent communication, leadership, and interpersonal skills with the ability to influence technical direction and drive complex initiatives across multiple teams.

RAKUTEN SHUGI PRINCIPLES: Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.

• Always improve, always advance. Only be satisfied with complete success - Kaizen.

• Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.

• Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.

• Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.

• Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151586549

Similar Jobs

Bengaluru, India

Skills:

TensorflowGitPytorchMLopsDockerFlaskFastAPIPythonKubernetesAWSAirflowvector databasesembedding modelsSagemakerStreamlit

Beware of Scammers

We don’t charge money for job offers