

Search by job, company or skills

Responsibilities
Manage and operate centralized monitoring and observability platforms across applications, databases, infrastructure, networks, and cloud environments to ensure 24/7 service availability.
Monitor system health using metrics, logs, and alerts proactively identify anomalies, performance issues, and service degradation.
Perform alert triage, impact assessment, and incident coordination, escalating issues to the appropriate technical teams to meet SLA requirements.
Design and enhance monitoring strategies, dashboards, alerting frameworks, and observability standards to improve service visibility and reduce alert noise.
Support major incident management by providing diagnostics, cross-team coordination, and driving service reliability improvements through trend analysis and root cause identification.
Monitor and optimize cloud and infrastructure costs, implementing tagging, budgeting, cost allocation, and identifying opportunities for cost savings.
Develop operational and cost reports, dashboards, and forecasts to support service management, leadership, and operational decision-making.
Drive continuous improvement by expanding monitoring coverage, automating observability processes, maintaining documentation, and supporting after-hours operational activities.
Requirements
3-5 years of experience in IT operations, NOC, service assurance, system monitoring, or cloud/infrastructure operations.
Hands-on experience with monitoring and observability tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or equivalent platforms.
Strong understanding of hybrid infrastructure (on-premises and AWS), system/network monitoring, application performance, and log/metric analysis.
Experience with AWS cost management, including Cost Explorer, budgeting, tagging strategies, and cloud cost optimization practices.
Familiarity with ITIL processes (Incident, Problem, and Change Management) AWS Associate-level certification or AWS FinOps Certified Practitioner is preferred.
Willingness to support after-hours operations, including deployments, maintenance, and incident response.
GMP Recruitment Services (S) Pte Ltd | EA Licence: 09C3051 | VO UYEN AI LINH | Registration No: R22109232
Job ID: 153750759
Skills:
Server Troubleshooting, Computer Engineering, Backup Solutions, Administration, Infrastructure Monitoring, Analytical and Problem-Solving Skills, Configuration Solutions, Infrastructure Automation, out-of-hours support, System Health Checks, Virtualization Platform, Infrastructure Deployment, Able To Work Independently, Collaborate With Internal Team, infrastructure Incident Management, server knowledge
Skills:
cohesity , San, Avamar, Vmware Vsphere, Veeam, Vlans, Gcp, Isilon, Data Domain, Azure, AWS, PowerStore, oci, Rubrik, VCF, Dell NetWorker
Skills:
cohesity , San, Vmware Vsphere, Avamar, Veeam, Vlans, Gcp, Isilon, Data Domain, Azure, AWS, PowerStore, oci, Rubrik, VCF, Dell NetWorker
Skills:
VMware, Backup, Windows Server, Logging, Dns, Firewalls, Red Hat Linux, routing, Nutanix, Load Balancing, Azure, AWS, Hyper-V, Monitoring, alerting, enterprise storage, Segmentation, observability
Skills:
VMware, Windows Server, Cisco Networking, Prometheus, Grafana, Backup Recovery, Openshift, Rhel, Kubernetes, FortiGate firewalls, infrastructure monitoring, Dell servers