Senior AI/ML Engineer
Location: Hyderabad
Experience: 6+ Years
Employment Type: Full-Time
Job Summary
We are looking for a highly skilled Senior AI/ML Engineer to design, build, deploy, and scale enterprise-grade Machine Learning and Artificial Intelligence solutions that deliver measurable business value.
The ideal candidate will have strong hands-on expertise in Machine Learning, MLOps, cloud platforms, ML model deployment, scalable ML pipelines, feature engineering, model monitoring, and software engineering practices. The candidate will work closely with Data Scientists, Data Engineers, Platform Engineers, SRE, DevOps, and Product teams to take AI/ML solutions from experimentation to reliable production systems.
This role requires a strong engineering mindset with the ability to build scalable, secure, observable, and maintainable ML solutions across the complete machine learning lifecycle.
Key Responsibilities
Machine Learning & AI Solution Development
- Design, develop, deploy, and scale machine learning and AI solutions for enterprise use cases.
- Translate business and product requirements into scalable AI/ML solutions.
- Work closely with Data Scientists to productionize machine learning models.
- Select appropriate ML algorithms, frameworks, and architectures based on business requirements.
- Develop reusable components and frameworks for AI/ML applications.
- Optimize models and ML applications for accuracy, performance, scalability, and cost.
ML Model Deployment & Productionization
- Deploy machine learning models into production environments across cloud and enterprise platforms.
- Build scalable real-time and batch inference solutions.
- Develop model-serving APIs and microservices.
- Implement model versioning, release management, rollback, and deployment strategies.
- Support production ML models and troubleshoot performance or reliability issues.
- Ensure models meet production SLAs for availability, latency, and throughput.
MLOps
- Design and implement end-to-end MLOps pipelines for model development, validation, deployment, and monitoring.
- Implement automated CI/CD pipelines for ML models and applications.
- Establish repeatable and reproducible ML workflows.
- Automate model training, testing, validation, deployment, and retraining processes.
- Implement model registry and experiment tracking solutions.
- Establish best practices for model lifecycle management and governance.
ML Pipeline Development
- Design scalable and reliable data and ML pipelines.
- Build automated workflows for data preparation, feature engineering, model training, and inference.
- Integrate data pipelines with ML workflows.
- Implement data validation and quality checks.
- Optimize pipelines for large-scale datasets and distributed processing.
- Support both batch and real-time ML processing requirements.
Feature Engineering & Feature Management
- Design and implement scalable feature engineering frameworks.
- Develop reusable feature transformation pipelines.
- Ensure consistency of features between training and production environments.
- Implement feature versioning and lifecycle management.
- Work with Data Engineering teams to build reliable feature pipelines.
- Exposure to Feature Store technologies is preferred.
Model Monitoring & Observability
- Design and implement comprehensive monitoring solutions for production ML models.
- Monitor:
- Model accuracy
- Data quality
- Data drift
- Concept drift
- Model performance
- Prediction latency
- Infrastructure health
- Develop dashboards, alerts, and operational metrics.
- Implement observability across ML pipelines and model-serving infrastructure.
- Identify model degradation and implement appropriate remediation strategies.
Cloud & Infrastructure
- Design and deploy AI/ML solutions on cloud platforms such as AWS, Azure, or GCP.
- Work with cloud-native services for ML, compute, storage, networking, and monitoring.
- Build scalable and highly available ML infrastructure.
- Implement secure cloud architectures for AI/ML workloads.
- Optimize cloud infrastructure and ML workloads for performance and cost.
- Exposure to Infrastructure as Code tools such as Terraform or CloudFormation is preferred.
Software Engineering
- Develop production-quality code primarily using Python.
- Build scalable APIs and microservices for AI/ML applications.
- Follow clean coding, design, testing, and documentation practices.
- Perform code reviews and contribute to engineering standards.
- Develop reusable libraries and components.
- Troubleshoot application, infrastructure, and ML pipeline issues.
CI/CD & DevOps
- Design and maintain CI/CD pipelines for ML applications and models.
- Automate build, test, validation, and deployment processes.
- Integrate ML workflows with enterprise DevOps practices.
- Implement automated quality and security checks.
- Work with Git-based source control and modern DevOps tools.
Model Governance & Security
- Implement appropriate model governance and lifecycle management processes.
- Maintain documentation for models, datasets, experiments, and deployments.
- Ensure ML solutions comply with enterprise security and privacy standards.
- Implement access controls and secure handling of sensitive data.
- Support auditability, traceability, and reproducibility of ML solutions.
- Contribute to responsible AI and model risk management practices.
Cross-Functional Collaboration
- Collaborate with:
- Data Scientists
- Data Engineers
- Platform Engineers
- SRE Teams
- DevOps Engineers
- Cloud Architects
- Product Managers
- Business Stakeholders
- Provide technical guidance on productionizing ML models.
- Participate in architecture and design discussions.
- Mentor junior AI/ML engineers and promote engineering best practices.
- Communicate complex technical concepts clearly to technical and business stakeholders.
Required Technical Skills
Machine Learning
- Strong understanding of machine learning concepts and algorithms.
- Experience with supervised and unsupervised learning.
- Knowledge of:
- Regression
- Classification
- Clustering
- Recommendation Systems
- Feature Engineering
- Model Evaluation
- Experience with ML frameworks such as:
- Scikit-learn
- XGBoost
- TensorFlow
- PyTorch
Programming
- Strong proficiency in Python.
- Good understanding of object-oriented programming and software engineering principles.
- Strong SQL knowledge.
- Experience developing REST APIs and microservices.
MLOps
- Hands-on experience with end-to-end ML lifecycle management.
- Experience with one or more:
- MLflow
- Kubeflow
- AWS SageMaker
- Azure ML
- Google Vertex AI
- Experience with experiment tracking and model registry.
- Model deployment and monitoring experience.
Cloud
Strong experience with at least one major cloud platform:
- AWS
- Microsoft Azure
- Google Cloud Platform
AWS ML experience such as Amazon SageMaker, S3, Lambda, ECS/EKS, IAM, and CloudWatch is highly desirable.
Data Engineering
- Data ingestion and transformation.
- ETL/ELT pipelines.
- Feature engineering pipelines.
- Batch and real-time processing.
- Experience with Apache Spark/PySpark is preferred.
- Familiarity with data lakes and data warehouses.
DevOps & Containerization
- Docker
- Kubernetes
- Git
- CI/CD
- Jenkins / GitHub Actions / GitLab CI
- Terraform or CloudFormation
Monitoring & Observability
- Prometheus
- Grafana
- Datadog
- CloudWatch
- Logging and alerting frameworks
- Model performance monitoring
Preferred / Good-to-Have Skills
- Experience with Generative AI and LLM-based applications.
- Knowledge of RAG architectures and vector databases.
- Exposure to LangChain or LangGraph.
- Experience with AI agents and agentic workflows.
- Experience with Kafka or other event-streaming platforms.
- Knowledge of Feature Stores.
- Experience with model explainability and Responsible AI.
- Exposure to AI/ML solutions within Banking, Financial Services, Healthcare, or other regulated industries.
Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, Data Science, Artificial Intelligence, or a related technical field.
- 6+ years of relevant professional experience in AI/ML Engineering, Machine Learning Engineering, Data Science Engineering, or related software engineering roles.
- Demonstrated experience taking ML solutions from development/experimentation through production deployment.
- Strong understanding of software development lifecycle and Agile methodologies.
Soft Skills
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
- Strong ownership and accountability.
- Ability to work independently and across multiple technical teams.
- Ability to manage priorities in a fast-paced environment.
- Strong mentoring and technical leadership capabilities.
- Passion for learning and adopting emerging AI/ML technologies.