

Search by job, company or skills

Key Responsibilities
Production Support & Incident Management
. Act as the primary Ll /L2 support contact for digital platforms, e-commerce systems, websites, and customer-facing services.
. Monitor incident queues, service requests, alerts, and support tickets, ensuring adherence to SLAs and operational procedures.
. Lead incident triage, troubleshooting, escalation, and resolution activities.
. Perform impact assessment and coordinate with relevant stakeholders during service disruptions.
. Support incident management activities and facilitate communication during critical outages.
. Conduct post-incident reviews and root cause analysis (RCA) to prevent recurrence.
. Develop and maintain operational runbooks, support procedures, and knowledge base articles.
System Monitoring & Reliability
. Monitor application, infrastructure, and business service health using observability and monitoring tools.
. Analyse system performance, availability, error trends, and capacity utilization.
. Configure and tune alerts to reduce noise and improve operational visibility.
. Collaborate with engineering teams to improve system reliability and operational resilience.
. Support Site Reliability Engineering (SRE)practices, including reliability metrics, incident reduction, and service availability improvements.
Cloud & Infrastructure Support
. Provide operational support for cloud-hosted applications and infrastructure, primarily on AWS.
. Perform first-level troubleshooting on:
o Compute services (EC2, ECS, Lambda)
o Networking
o Load Balancers
o CDN services
o Storage services
. Investigate infrastructure-related issues affecting application performance or availability.
. Support deployment verification and post-release monitoring activities.
Application & Integration Support
. Troubleshoot application issues across web, mobile, APIs, and middleware platforms.
. Analyse application logs, monitoring data, and system traces to identify root causes.
. Support integrations with external systems, partners, payment gateways, and other enterprise platforms.
. Work closely with L3 to reproduce issues and validate fixes.
. Support release and deployment activities, including late-night and weekend implementations when required.
Continuous Improvement
. Identify recurring incidents and operational inefficiencies.
. Drive automation opportunities to reduce manual effort and repetitive support activities.
. Recommend improvements to monitoring, alerting, deployment processes, and support workflows.
. Contribute to operational excellence initiatives and service reliability improvements.
Stakeholder &Vendor Management
. Collaborate with internal teams, external vendors, and partners across different geographies and time zones.
. Communicate effectively with technical and non-technical stakeholders.
. Provide timely updates during incidents and service disruptions.
. Participate in operational reviews, governance meetings, and service improvement discussions.
Required Skills & Experience Technical Skills
. 5+ years of experience in Application Support, Production Support, Technical Operations, TechOps, or related roles.
. Strong knowledge of AWS cloud services and operational support.
. Good understanding of cloud-native application architecture.
. Experience supporting:
o Digital platforms
o E-commerce systems
o Customer-facing web applications
o Mobile applications
. Knowledge of CDN technologies and content delivery architecture.
. Experience with monitoring and observability platforms such as:
o Datadog
o New Relic
o Dynatrace
o AppDynamics
o Grafana
o CloudWatch
. Familiarity with APM (Application Performance Monitoring) concepts.
. Understanding of SRE principles and operational best practices.
. Strong knowledge of:
o APIs
o Microservices
o Web services
o System integrations
o Authentication and authorization flows
. Experience using ticketing and ITSM platforms(Jira Service Management, ServiceNow, etc.).
. Understanding of Incident, Problem, and Change Management processes.
. Familiarity with log analysis tools such as Elasticsearch, Kibana, Splunk, or Cloud WatchLogs.
Soft Skills
. Strong troubleshooting and analytical thinking abilities.
. Naturally curious and investigative, with adesire to understand the full context behind incidents and operational events.
. Excellent problem-solving and root cause analysis skills.
. Strong ownership mindset and accountability.
. Ability to work independently in a fast-paced operational environment.
. Good communication and stakeholder management skills.
. Ability to remain calm and methodical during high-severity incidents.
. Strong documentation and knowledge-sharing practices.
. Continuous improvement mindset with a focus on automation and operational efficiency.
Good to Have Skills
. Knowledge of WeChat Mini Program ecosystem and integrations.
. Experience supporting SAP Commerce, Adobe Experience Manager (AEM), Magento, Shopify, or similar e-commerce platforms.
. Basic scripting skills (Python, Shell, Bash, PowerShell).
. Experience with API Gateway and event-driven architectures.
. AWS Certifications (Cloud Practitioner, Solutions Architect Associate, SysOps Administrator).
. ITIL Foundation certification.
. Experience supporting payment gateways and digital commerce ecosystems.
Working Conditions
. Participate in a 24x7 support and on-call rotation model.
. Support late-night, weekend, and public holiday deployments where required.
. Work closely with internal teams, vendors, and stakeholders across multiple time zones.
. Respond to critical production incidents outside business hours when necessary.
Job ID: 152322503