Job Summary:
As an ICT Infrastructure Engineer in the Authority's Data Analytics & Engineering (DAE) Division, Production Engineering (PE) team, you will manage multiple priorities in a fast-paced environment while supporting stakeholder needs and responding to change. Specifically, you are expected to:
- Partner with DAE Data Science and Product Management teams to productionise and host Data Science and AI products on PE-managed on-premises and cloud AI platforms.
- Work with Technology Group divisions-including Platform Architecture and Engineering, Cybersecurity, and Data & Collaboration Platforms-to define requirements and integrate scalable, secure, and compliant AI platforms into the Authority's enterprise environment.
Key Responsibilities:
- Deliver projects within the assigned practice to meet platform availability, reliability and resilience requirements.
- Translate business requirements into platform infrastructure designs.
- Work with service providers to design, deploy and configure resilient, available and secure AI infrastructure.
- Ensure compliance with the Authority's Enterprise Architecture and Security Standards, as well as IM8 and AI Policies.
- Act as the point-of-contact for new AI product initiatives, supporting system design, integration, acceptance and performance testing.
- Manage AI platform operations and address queries from Data Scientists and Product Managers.
- Improve operational runbooks, SOPs and adhere to enterprise and compliance mandates.
- Continually improve AI platform experience by implementing new features and enhancements.
- Continually simplify product development and operations through automation and building reusable services.
- Assist in AI Platform design review and improvement of AI, OS, container and database security standards.
- Onboard hosting environments by designing, setting up, configuring, implementing, testing and commissioning the required infrastructure.
- Manage system architecture, resource planning, platform performance, and hosting across servers, operating systems, and containers, excluding hypervisor and storage layers.
What we are looking for:
- Strong written and verbal communication skills, with the ability to pitch ideas, influence stakeholders and balance strategic perspective with business needs and challenges.
- Ability to assess project and compliance requirements and develop practical, sustainable solutions.
- Hands-on experience building, securing and operating enterprise Data Science and AI platforms on AWS or in on-premises environments.
- Hands-on experience building and operating DevSecOps, MLOps, and/or LLMOps pipelines.
- Hands-on experience with AI and container services, LLM inference and Linux preferred.
- Hands-on experience with Infrastructure-as-Code, automation and Python tech stack preferred.
- Experience working with Identity & Access Management services (e.g. Entra ID, AWS Cognito and AWS IAM), and Monitoring & Observability services (e.g. CloudWatch, Splunk and Elastic) are advantageous.
- Knowledge of PostgreSQL, vector databases, object storage, and the processing of structured and unstructured datasets is advantageous.
- Good understanding of AI, application and infrastructure security risks and controls.
TG-DAE-REQ126-3 / TG-DAE-REQ126-2 / TG-DAE-REQ126-1