Software Engineer, Cloud
JOB INFO
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
COMPANYOllama
About the role You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.
What you'll do
- Build and scale the inference platform that serves every request from ollama.com.
- Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.
- Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.
- Build the reliability, observability, and cost controls for our team and customers.
You may be a fit if
- You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.
- You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.
- You've worked with Kubernetes, GPU scheduling, or inference infrastructure.
- You think in terms of reliability, SLOs, and honest capacity planning.
- Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.