About the Role
We are looking for an experienced Senior Full-Stack Engineer to design and deliver a high-scale, AI-enabled platform supporting the digitisation and processing of national examination scripts.
This is a hands-on senior engineering role spanning full-stack development, AI-integrated pipelines, desktop applications, hardware integration, and system architecture. You will work on a platform that combines web applications, Windows-based scan-room software, physical document scanners, computer vision, OCR, and LLM-powered workflows.
The platform processes millions of examination pages within demanding operational windows, making reliability, correctness, scalability, and recoverability critical engineering requirements.
You will take ownership of key architectural decisions, build production systems hands-on, establish strong engineering practices, and provide technical guidance to other engineers in the team.
Key Responsibilities
- Design, develop, and maintain full-stack applications and services supporting large-scale document scanning and examination processing workflows.
- Architect and implement AI-integrated processing pipelines using LLMs, computer vision, OCR, and related AI services.
- Build robust orchestration and validation layers around AI components, including retry handling, output validation, error classification, prompt/version management, and observability.
- Define appropriate preconditions, postconditions, validation rules, and system invariants to ensure AI-generated outputs are safe and reliable for downstream processing.
- Design systems that can handle high-volume processing, including peak examination periods involving large numbers of scanned pages and concurrent processing jobs.
- Build resilient workflows capable of handling partial failures, service interruptions, rate limits, hardware failures, and interrupted processing without compromising overall system integrity.
- Develop and maintain backend services, APIs, databases, job queues, and asynchronous processing workflows.
- Develop and enhance frontend applications, including interactive interfaces for document and template configuration.
- Build and maintain Windows desktop applications used by scan-room operators, including integration with physical document scanners and vendor SDKs.
- Troubleshoot issues across application, operating system, hardware, driver, network, and cloud service boundaries.
- Design appropriate boundaries between deterministic processing and AI/agentic workflows, considering correctness, performance, latency, and cost.
- Implement monitoring and observability for AI pipelines, including input/output distributions, failure rates, latency, model behaviour, and operational cost.
- Establish and promote engineering practices including test-driven development, automated testing, typed interfaces, CI/CD quality gates, logging, monitoring, and operational runbooks.
- Lead technical and architectural discussions and document key decisions through architecture documentation and decision records.
- Conduct code reviews and mentor junior and mid-level engineers, helping raise engineering standards across the team.
- Work closely with product, operations, infrastructure, and other engineering stakeholders to ensure the platform remains production-ready during critical operational periods.
Requirements
- Degree in Computer Science, Software Engineering, Information Systems, or a related discipline.
- Strong hands-on experience in full-stack software engineering, with experience designing and operating complex production systems.
- Strong backend engineering capabilities and experience designing APIs, asynchronous workflows, job queues, databases, and distributed processing systems.
- Strong frontend development experience with modern web technologies such as React, TypeScript/JavaScript, or equivalent.
- Strong understanding of software architecture, algorithmic complexity, concurrency, asynchronous programming, state management, and distributed-system failure modes.
- Experience designing systems for high availability, scalability, fault tolerance, and operational recoverability.
- Production experience integrating LLMs, computer vision, OCR, or other AI/ML services into software applications.
- Ability to design robust AI integration layers covering validation, retries, failure handling, model/prompt versioning, monitoring, and evaluation.
- Understanding of AI-specific operational considerations including model drift, latency variability, token/context constraints, prompt injection, output uncertainty, and cost management.
- Experience defining measurable evaluation criteria and automated checks for AI-generated outputs.
- Strong automated testing practices across unit, integration, end-to-end, and system-level testing.
- Experience with CI/CD, application monitoring, structured logging, observability, and production troubleshooting.
- Strong problem-solving skills and ability to diagnose complex issues spanning multiple components or system boundaries.
- Experience providing technical leadership, architecture guidance, code reviews, and mentoring to other engineers.
Good to Have
- Experience building multi-step or agentic AI workflows, including orchestrator-worker, generator-critic, parallel processing, or similar patterns.
- Experience with MCP or other tool orchestration and permission frameworks.
- Experience developing and deploying desktop applications, such as Electron-based, native Windows, or equivalent applications.
- Experience integrating software with physical devices, scanners, peripherals, vendor SDKs, or hardware drivers.
- Experience with document processing, OCR, image processing, computer vision, or OMR systems.
- Experience operating systems with high-volume batch or document-processing workloads.
- Experience working within government, regulated, security-sensitive, or other high-assurance environments.
- Experience producing architecture decision records, operational runbooks, technical documentation, and onboarding materials.
What We Value
We are looking for an engineer who treats AI as a production system component rather than simply an API call. You should be comfortable reasoning about what happens when a model returns an incorrect or malformed result, an external AI service becomes unavailable, a processing job fails halfway through, or a physical scanner stops responding during an operational run.
You should also be comfortable deciding when AI should not be used, and when a deterministic implementation provides better correctness, latency, cost, or operational reliability.
Most importantly, you should be someone who remains hands-on while providing technical leadership — able to move between architecture, implementation, debugging, code review, and mentoring as required.