Search by job, company or skills

Production Support Engineer

5-7 Years
SGD 13,000 - 16,000 per month
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

Salesforce is partnering with Alibaba Cloud to offer its core products in China, helping support existing multinational customers who want to expand their presence in the Greater China Region (GCR) and need Salesforce to run in China to meet the China regulatory guidance.

The Hyperforce on Alibaba Cloud Production Support Engineering team is an Organization within Security Engineering, with an exciting mission to bootstrap adoption of the industry's leading-edge SRE principles and best practices at Salesforce. We are looking for experienced Software Engineers/DevOps Engineers with a strong automation mindset, a proven track record of delivering automation solutions with hands-on experience troubleshooting Infrastructure and Java applications to join this growing team. Working closely with counterparts in the Infrastructure and Engineering organizations, this Hyperforce on Alibaba Cloud Production Support Engineering group owns the reliable delivery of service to Salesforce engineering teams and customers running on Alibaba cloud infrastructure. This organization provides round-the-clock, follow-the-sun situational awareness and leadership in the swift resolution of any service-impacting issues, driving customer success.

As a member of the team, you will be responsible for detecting, investigating and resolving system failures and complex outages across infrastructure and java applications, including creation of the observability tooling necessary for your success. This objective is met by monitoring the services, reacting to problems, proactively addressing issues before they affect performance or availability, and working with Engineering teams to define service level objectives and improving service design and implementation to increase reliability through closed-loop feedback. This engineer should be comfortable working through ambiguity, coordinating with multiple engineering teams.

Hyperforce on Alibaba Production Support Engineer focuses more on proactive automation, and targets 50%+ time spent on improving service design for reliability, extending monitoring and operational automation, driving self-healing and resiliency initiatives and game day exercises. The incumbent in this role would demonstrate a strong focus on tactical operations, as well as large-scale production engineering and orchestration. Core Competencies focus on technical troubleshooting and investigation, where engineers serve as experts leading the resolution of complex customer issues and production trends. The role involves bridging gaps by acting as a facilitator between development, infrastructure, and support teams to eliminate finger-pointing during cross-product incidents and drive progress. Furthermore, the team champions customer advocacy by being the voice of the customer within Engineering, influencing bug prioritization and service design based on real-world impact. Finally, a commitment to operational excellence ensures the delivery of comprehensive production health checks and training programs to improve the supportability and reliability of Salesforce products.

The successful candidate will demonstrate strong ownership, initiative, and accountability. They will be able to prioritise work independently, communicate progress and risks clearly, and take an idea from initial investigation through design, implementation, rollout, and measurable operational improvement.

Minimum qualifications:

  • BS or MS in Computer Science or a related technical field involving systems engineering

  • Experience will be evaluated based on alignment to the core competencies for the role (e.g. extracurricular leadership roles, military experience, volunteer work, etc.).

  • 5+ years infrastructure and applications systems engineering experience in enterprise-scale Internet services. Experience in analyzing and troubleshooting systems using logging, distributed tracing, stack traces, and debuggers

  • 5+ years experience configuring and managing any of the Public Clouds using CLI/SDKs and automation (Alibaba or AWS preferred)

  • 5+ years experience in at least one of the following languages: Java, Python, Go. Ability to pick up new languages

  • Experience in Unix/Linux environments with good understanding of operating systems internals (e.g., filesystems, system calls)

  • Working knowledge of the TCP/IP stack, routing and load balancing technologies

  • Working knowledge of design principles of monitoring and alerting systems

  • Ability to operate in a high-pressure environment, troubleshoot complex issues quickly, and successfully handle multiple priorities

  • Systematic problem-solving approach, coupled with a strong sense of ownership and drive

  • Incident management - Act in key support roles during major incidents e.g. Sev0, Sev1. Also, participate in the technical review of the incident for problem management

  • Experience leading complex technical troubleshooting and investigation of critical customer issues and production trends.

  • Proven ability to act as a facilitator between development, infrastructure, and support teams to resolve cross-product incidents.

  • Strong communication skills in English and Mandarin is required to work with customers from the Greater China Region.

  • Perform Schedule Change Request

  • Experience building with Claude Code or similar AI development tools

Preferred qualifications:

  • A good understanding and practice in large-scale distributed systems

  • Experience in designing and deploying high performance production services with extensive monitoring and logging practices

  • CI/CD automation experience, including understanding of key open source technologies like Jenkins, Spinnaker, and Docker

  • Experience defining immutable infrastructure via Terraform/Cloud Formation or other approaches across large footprints and distributed teams

  • Experience with on-call rotation, leading incident response and no-blame postmortem analysis

  • Ability to debug, optimize code, and automate routine tasks

  • Customer/Partner Facing Experience

  • Strong track record of customer advocacy, influencing service design and bug prioritization based on real-world customer impact.

  • Leadership experience in major incident management (Sev0/Sev1) and driving operational excellence through comprehensive health checks.

  • Have release/build experience

More Info

Job Type:
Industry:
Employment Type:

Job ID: 151721203