Your role
As a Sr. AI Engineer: Systems, you’ll serve as an embedded senior back-end engineer on our Speech Team, owning the production systems that turn speech models and third-party capabilities into reliable, scalable experiences for Dialpad’s AI voice agents. You’ll work at the intersection of speech, ML infrastructure, and product engineering: productionizing models, enabling self-hosted inference, integrating external APIs, and building the operational foundations required for strong uptime and latency SLAs. You’ll partner closely with the MLOps (Inference) team while bringing deep ownership of the speech domain, helping the team move quickly from promising model or vendor capability to safe, observable, and cost-effective production. This role offers broad technical influence and the opportunity to shape how Dialpad operates real-time speech systems at scale.
This position reports to our Senior Manager, AI Speech, and offers the opportunity to be based in our Canada Hub locations.
What you’ll do
- Productionization & Service Ownership: Own the path from speech model or third-party capability to production, building the APIs, services, deployment workflows, and integration layers that make it safe and easy for the Speech Team to ship improvements.
- Self-Hosted Inference & Scaling: Productionize and operate self-hosted speech models, optimizing serving architecture, resource utilization, concurrency, autoscaling, and cost so they can meet the demands of real-time voice agents.
- Third-Party APIs & Provider Resilience: Integrate and maintain third-party speech APIs behind durable abstractions, with clear failover, capacity planning, version management, and vendor-performance monitoring.
- Reliability, SLOs & Observability: Build the monitoring, alerting, dashboards, health checks, and incident-response practices needed to meet uptime, latency, and quality SLAs for customer-facing speech systems.
- Release & Evaluation Infrastructure: Partner with Speech and MLOps engineers to enable shadow traffic, staged rollouts, model and artifact versioning, rollback-safe releases, and candidate-versus-incumbent comparisons.
- Cross-Functional Leadership & Mentorship: Work closely with the MLOps (Inference) team and partner teams across speech, platform, telephony, and product to set technical direction, mentor engineers, and turn model advances into reliable production impact.
Skills you’ll bring
- Systems & Backend Engineering: Strong software engineering fundamentals and proficiency in Python, plus experience designing maintainable APIs, services, and integration layers. We’re open to candidates who are strongest in backend/platform engineering or who have grown from ML into systems.
- Production ML & Streaming: 5+ years of experience building or operating production software, including ML-backed systems, real-time services, speech applications, streaming media, or other latency-sensitive systems.
- Model Serving & Inference: Hands-on experience deploying, scaling, and troubleshooting ML models in production, including model serving, inference optimization, resource management, and safe model and version rollouts.
- Cloud & Distributed Systems: Experience with cloud infrastructure and distributed systems, as well as familiarity with containers, orchestration, service networking, CI/CD, and GCP.
- Reliability & Operations: Strong understanding of observability, alerting, incident response, capacity planning, and availability and latency SLAs for customer-facing systems.
- Cross-Functional Technical Leadership: Demonstrated ability to work closely with ML scientists, MLOps and inference engineers, and product teams, mentor teammates, and make pragmatic trade-offs across quality, reliability, latency, scale, and cost.