Site Reliability Engineer

JOB INFO

Apply

Apply for this job directly on SHORTList.

Referral

Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.

COMPANYHelsing

Much of our work takes place in high-security on-premise environments, and we are looking for Site Reliability Engineer to support our high security environments. Your role as a Site Reliability Engineer will be to design, implement, and manage our on-premise Kubernetes infrastructure. We are looking for engineers with a strong work ethic and prioritisation skills. We value team players who communicate clearly, share knowledge generously, and collaborate effectively to move their team — and our mission—forward.

As a SRE, you will design and build cloud-native infrastructure platforms on-premises, focusing on Kubernetes-based solutions that enable our development teams to operate services at scale. You will create robust observability frameworks using Grafana, Prometheus, and distributed tracing to ensure system reliability and performance. You will architect and implement secure, multi-tenant Kubernetes clusters with strong access controls, policy-as-code governance, and zero-trust networking between red and black network domains. You will develop operators and controllers to automate infrastructure provisioning and compliance. You will build and maintain MLOps platforms enabling AI researchers to deploy, monitor, and scale machine learning models in production. You will collaborate closely with our Security teams to implement supply chain security, container scanning, and runtime protection across our cloud-native stack.