Principal Site Reliability Engineer, Infrastructure & Platform
f5 networks singapore pte ltd- Posted 8 hours ago
- Be among the first 10 applicants
Job Description
Responsibilities:
Infrastructure Automation & Configuration Management
- Author, maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment
- Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions
- Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration
- Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets
Compute & Virtualization
- Deploy and manage Proxmox VE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication
- Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, and Proxmox API automation
- Manage physical server provisioning end-to-end via HPE iLO (firmware updates, SPP deployment, OS installation via virtual media)
Container & Kubernetes Platforms
- Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management
- Operate Docker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring
- Maintain container image pipelines and registry infrastructure (Azure Container Registry or AWS ECR)
Cloud Platforms
- Engineer and maintain infrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)
- Apply cloud cost awareness, security best practices, and IaC principles (IAM, security groups, networking, storage) across AWS and Azure environments
Networking & Core Services
- Operate and troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)
- Maintain directory services (OpenLDAP master-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication
- Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets
Observability & Security
- Maintain and extend monitoring infrastructure (Prometheus, Observium) across a global fleet including SNMP polling, metrics collection, and alerting
- Manage centralised log aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs
- Operate runtime security tooling (Falco) and file integrity monitoring (AIDE) in production environments
- Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews
Reliability & Incident Response
- Participate in a 24x7 on-call rotation, responding to and leading production incident resolution
- Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements
- Define and track SLOs/SLIs for critical infrastructure services
- Identify and address single points of failure design and implement HA improvements
Requirements:
Strong Linux systems administration skills (RHEL/CentOS preferred) including systemd, networking, storage, kernel tuning, and package management
Proficiency with Ansible (or similar tool) for large-scale configuration management, including role design, inventory management, and CI/CD integration
Hands-on experience with at least one hypervisor platform, preferably Proxmox VE, Harvester (Kubevirt) or similar (VMware vSphere, KVM)
Production experience operating on-premise Kubernetes clusters (rke2, k3s, etc)
Practical AWS or Azure experience including compute, networking (VPC/VNet, security groups, DNS), IAM, and managed services
Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, and firewall rule management
Experience with secrets management platforms (HashiCorp Vault or equivalent)
Familiarity with PCI-DSS requirements as they apply to infrastructure -- hardening standards (CIS benchmarks), audit logging, access control
Experience writing and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)
Demonstrable on-call experience and comfort leading incident response in a global production environment
- 7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment
More Info
Key Skills
Managed Network Services
PCI Standards
Azure Key Vault
Tools Software
Virtualization Management
Security Platform
