

Search by job, company or skills
Location: Singapore
Working Hours: Sunday to Thursday | 11:00 AM to 7:00 PM
This is not an internal IT role.
This is not a helpdesk position.
This is a vendor-grade technical leadership role operating at OEM engineering standards, supporting mission-critical server and storage environments across multiple regions.
ABOUT THE ROLE
Maintech is looking for a hands-on Level 3 Subject Matter Expert who can take full ownership of complex technical escalations, perform deep troubleshooting and root cause analysis, and provide final technical direction for critical customer issues.
You will act as the senior technical escalation point for incidents that cannot be resolved by L1 and L2 teams.
This role requires:
• Strong analytical and troubleshooting capability
• Technical ownership from investigation through resolution
• Confidence in making decisions during critical incidents
• Strong customer and cross-functional communication
• The ability to operate effectively in mission-critical environments
This is not a ticket coordination or pure people management role. The successful candidate must remain technically hands-on.
WHAT YOU WILL OWN
L3 Technical Escalation
• Take ownership of complex incidents escalated by L1 and L2 support teams
• Serve as the senior technical escalation point for high-impact, recurring and business-critical issues
• Assess technical impact, urgency, risk and resolution priorities
• Provide final technical direction and drive incidents through to closure
• Work with engineering teams, vendors, product teams and regional SMEs when required
Deep Troubleshooting and Root Cause Analysis
• Analyze system behavior, logs, error messages, configurations and performance data
• Identify the actual root cause instead of providing only temporary workarounds
• Validate findings through evidence, testing and technical analysis
• Develop corrective and preventive actions to reduce recurring incidents
• Clearly distinguish between service restoration and permanent resolution
Major Incident and Critical Issue Management
• Participate in major incident calls and technical war rooms
• Provide technical direction and recommendations during critical incidents
• Support rapid decision-making under time-sensitive and high-pressure conditions
• Communicate incident progress, technical risks and recovery plans to customers and internal stakeholders
• Complete RCA reports, incident reviews and follow-up improvement actions
Technical Solutions and Risk Assessment
• Recommend remediation, upgrade, migration or infrastructure improvement solutions
• Assess technical risks and potential business impact before changes are implemented
• Support complex deployments, migrations, upgrades and system integration activities
• Ensure proposed solutions comply with customer SLAs, security policies and operational procedures
• Provide technical recommendations that improve system stability, performance and reliability
Knowledge Transfer and Team Development
• Guide L1 and L2 engineers through complex technical issues
• Mentor engineers and improve the team's overall troubleshooting capability
• Develop troubleshooting guides, SOPs, technical runbooks and knowledge base documentation
• Conduct internal or customer-facing technical training
• Improve escalation procedures and standardize troubleshooting practices
WHAT WE ARE LOOKING FOR
• Strong hands-on experience in enterprise systems, infrastructure or a relevant technical domain
• Proven experience handling L3 escalations, complex troubleshooting or mission-critical incidents
• Ability to independently perform log analysis, root cause analysis, solution design and risk assessment
• Experience with incident, problem and change management processes
• Strong ownership and the ability to drive issues through to closure
• Experience supporting enterprise customers or large-scale environments
• Ability to clearly explain complex technical issues to customers and non-technical stakeholders
• Professional English communication skills for technical meetings, written reports and cross-border collaboration
• Willingness to remain technically hands-on rather than operating only as a people manager or coordinator
PREFERRED QUALIFICATIONS
• Previous experience as a Technical Lead, Lead Engineer, Escalation Engineer or Subject Matter Expert
• Experience supporting regional or global teams across multiple time zones
• Experience in financial services, government, telecommunications, manufacturing, data centers or other highly regulated environments
• Familiarity with ITIL, incident management, problem management and change management
• Relevant certifications in infrastructure, cloud, server, network, storage or cybersecurity
• Experience with automation, scripting, monitoring or reporting tools
• Experience creating SOPs, technical runbooks, knowledge bases or training materials
• Experience participating in migration, deployment, upgrade or infrastructure improvement projects
• Experience working directly with OEM engineering teams, product teams or global SMEs
THIS ROLE MAY NOT BE SUITABLE FOR YOU IF
• Your experience is mainly limited to monitoring, ticket assignment or incident coordination
• Complex issues are normally transferred to vendors without your continued technical ownership
• You have moved entirely into people management and no longer perform hands-on troubleshooting
• Your experience is limited to standard procedures and known issue playbooks
• You have not personally conducted root cause analysis or developed permanent corrective actions
• You prefer to escalate issues rather than remain accountable until final resolution
Job ID: 150610251
Skills:
application infrastructure , change management, Problem Management, Application Lifecycle Management, Servers, Incident Management, System Architecture, Memory, PAS-X, validation requirements, system capacity, Siemens Opcenter, IT-OT integration, industrial IoT platforms, ERP integrations, Hardware, data historians
Skills:
data warehouses , Agile Development Methodologies, Pyspark, Sql, Scalability, Devops, Performance, MLops, Databricks, Python, AI solution design, GenAI solution patterns, Relational Databases, ML engineering, Resilience
Skills:
operational support , production readiness , Prometheus, Grafana, Terraform, MySQL, AWS, Redis, Systems Administration, Gcp, Ansible, Incident Management, Cloud Infrastructure, automated deployment, site reliability practices, Disaster Recovery, CI CD, alerting, access management, technical tools, observability, cloud platforms, IT setup, Monitoring, centralized logging, Root Cause Analysis, GitLab CI CD, infrastructure as code