
Search by job, company or skills
Responsibilities
. Define and drive the organization's cloud and platform infrastructure strategy, architecture, standards, and multi-year roadmap, ensuring scalable, secure, resilient, and cost-efficient solutions.
. Serve as the L3/L4 technical authority and final escalation point for complex infrastructure, platform, and cross-domain incidents, leading root cause analysis (RCA) and implementing permanent resolutions.
. Design, develop, and maintain Infrastructure as Code (IaC), platform automation, and GitOps practices using technologies such as Terraform and Python/Go.
. Architect, implement, and continuously improve platform resilience, high availability, disaster recovery (DR), fault tolerance, and service reliability, including defining and maintaining SLOs and SLIs.
. Design, optimize, secure, and manage enterprise Kubernetes environments, including networking, security, lifecycle management, and platform operations.
. Establish, govern, and enforce platform engineering, security, infrastructure, and CI/CD standards, while mentoring engineers and promoting engineering best practices.
. Independently manage and resolve complex incidents, service requests, problems, and changes within agreed SLAs, ensuring accurate documentation, timely ticket updates, stakeholder communication, and appropriate escalations.
. Proactively identify, investigate, analyse, and resolve platform issues, leveraging advanced troubleshooting techniques, operational diagnostics, and root cause analysis to prevent recurrence.
. Collaborate with clients, stakeholders, cross-functional teams, and automation teams to deliver platform improvements, optimise operational efficiency, and automate routine tasks.
. Produce and maintain technical documentation, share knowledge, coach L1-L3 engineers, and contribute to quality assurance, operational excellence, and continuous service improvement.
. Lead or contribute to infrastructure projects, platform enhancements, disaster recovery implementation and testing, and other technology initiatives as required.
. Perform other related duties as assigned.
Requirements
. Bachelor's degree in Information Technology, Computer Science, or a related discipline (or equivalent practical experience).
. Strong expertise in virtualization, cloud infrastructure, storage, automation, and enterprise platform technologies.
. Hands-on experience with VMware, Azure, AWS, Ansible, GitLab, DevOps/SRE practices, containers, and enterprise tools such as Veeam, Rubrik, Splunk, CyberArk, Opswat, NVIDIA AI, and storage platforms.
. Relevant industry certifications are highly desirable, including VMware, Microsoft Azure, AWS, Veeam, and Rubrik certifications.
. Excellent understanding of IT change management, with experience planning, assessing risks, executing changes, and documenting mitigation plans.
. Strong communication and collaboration skills with cross-functional, multicultural teams and stakeholders.
. Proven ability to work effectively in a fast-paced, high-pressure environment while managing multiple priorities.
. Strong client-focused mindset with a commitment to delivering exceptional service and customer experience.
. Excellent planning, organizational, problem-solving, and active listening skills, with the flexibility to adapt to changing business needs.
Licence no: 12C6060
Job ID: 151327927