Why This Job is Featured on The SaaS Jobs
Why this Role is Featured on The SaaS Jobs
AI-native SaaS products increasingly depend on reliable inference as a core production workload, not a sidecar experiment. This AI Systems Engineer role stands out because it is centered on the operational layer that turns in house models into low latency services on cloud GPU infrastructure, with clear attention to rollout safety, traffic handling, and observability. The remit suggests a SaaS organisation moving beyond prototypes into repeatable platform patterns for shipping model backed capabilities.
For SaaS career development, the work maps to durable platform competencies that recur across modern vendors: building shared internal infrastructure, defining standards for packaging and versioning, and creating measurable performance and reliability baselines. Experience with benchmarking, canary style releases, and cost performance tradeoffs is especially portable in SaaS, where unit economics and service level expectations shape engineering priorities as much as model quality.
The section above is editorial commentary from The SaaS Jobs, provided to help SaaS professionals understand the role in a broader industry context.
Job Description
Your role
We are hiring AI Inference Platform Engineers to build the production systems that serve our in-house AI models at scale.
This role sits at the intersection of model development, high-performance runtime systems, and cloud infrastructure. You will help turn trained models and emerging AI capabilities into reliable, observable, low-latency production services running on NVIDIA GPUs in GCP.
This is not a research role, and it is not a generic MLOps or support role. It is an implementation-heavy systems engineering role focused on the machinery of inference: model serving, runtime optimization, GPU utilization, deployment safety, traffic management, benchmarking, and production reliability.
Our mission is to shorten the path from promising model capability to dependable production impact. We build the shared infrastructure, standards, and release pathways that allow models to move from candidate artifacts into scalable, rollback-safe inference services with clear performance, reliability, and cost characteristics.
This is a new team, so the systems and interfaces are still being shaped. You will help define how models are packaged, deployed, benchmarked, monitored, compared, and operated across environments. The work is practical, deeply technical, and closely tied to the company’s broader AI strategy. We are not building one-off demos; we are building the inference platform by which a growing AI organization can repeatedly and safely ship real model-backed products.
What you’ll do
- You will design, build, and improve the systems that connect AI capability development to production inference.
- Depending on your strengths, your work may include:
- Inference Serving & Runtime Systems: Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
- GPU Infrastructure & Utilization: Operate and optimize containerized workloads on Kubernetes/GCP, with a focus on efficient use of NVIDIA GPUs, memory, storage, and networking.
- Model Server Integration: Work with model-serving frameworks and runtimes such as vLLM, Triton, TGI, or similar systems, adapting them to internal deployment, observability, and release requirements.
- Traffic & Release Safety: Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services.
- Benchmarking & Evaluation Infrastructure: Build tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
- Artifact Lifecycle: Improve how model and capability artifacts are packaged, versioned, promoted, deployed, and rolled back across environments.
- Observability & Debuggability: Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting so engineers can understand model-serving behavior in production.
- Efficiency & Scale: Contribute to strategies that improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.
Skills you’ll bring
- Production Engineering Experience: 6+ years of professional software engineering experience, with a track record of shipping backend services, infrastructure systems, or production platforms that matter.
- Strong Software Fundamentals: Proficiency in writing maintainable production code in Python, Go, or another backend-oriented language, with strong debugging and systems-thinking skills.
- Inference or Systems Orientation: Experience building, operating, or optimizing high-throughput services, distributed systems, data/ML infrastructure, or runtime platforms where latency, reliability, and resource utilization matter.
- Kubernetes & Linux Fluency: Hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
- Operational Judgment: A strong instinct for reproducibility, observability, rollout safety, failure modes, and whole-system resilience.
- Performance Awareness: Comfort reasoning about bottlenecks across compute, memory, network, storage, batching, concurrency, and service-level objectives.
- Collaboration: Ability to work closely with model developers, product engineers, infrastructure teams, and technical leadership to turn evolving AI capabilities into reliable production systems.