Why This Job is Featured on The SaaS Jobs
This Lead Product Manager role sits in a distinctly SaaS-critical layer: the AI model and inference stack that multiple customer-facing products depend on. In a market where AI features are increasingly table stakes, the differentiator often becomes reliability, latency, and measurable quality in production. The listing signals a product function embedded directly in model lifecycle work rather than an API integration surface, which is still relatively rare in SaaS org design.
For a long-term SaaS career, the role builds durable instincts around operating AI as a product, including evaluation discipline, production monitoring, and explicit trade-offs between quality, cost, and real-time constraints. That combination translates across SaaS companies shipping AI capabilities, especially where internal platform teams must serve many downstream product teams and where roadmap credibility depends on observable system behavior.
This is best suited to professionals who prefer decision-making grounded in data, incident learnings, and written narratives over ceremony-heavy delivery rituals. It will fit someone who wants to stay close to technical truth while still owning product calls, particularly those moving from applied ML or research into product leadership and seeking accountability that matches authority.
The section above is editorial commentary from The SaaS Jobs, provided to help SaaS professionals understand the role in a broader industry context.
Job Description
Your role
Our AI organization builds, runs, and hosts the models behind our products: custom SLMs, our ASR stack, and the inference infrastructure that serves them in real time. Products like our voice agents and agentic runtime are joint efforts with product engineering — but the models they run on are built here, and this role owns product for exactly that layer. It's a different job from the rest of platform and product engineering: research-driven, eval-heavy, and closer to training data and model behavior than to sprint boards. The standard PM toolkit doesn't cover it. The day-to-day runs on eval reports, latency budgets, and knowing whether a failure is a model problem, a serving problem, or a prompt problem — and without that fluency, even an excellent PM ends up coordinating from the outside instead of deciding from the inside. This is not a role you can do at the API-orchestration level. You need to know how models are built and run, ideally because you've built them.
We're hiring someone who won't have that problem. You've been on the other side of the table — as an AI researcher, applied scientist, or ML/AI engineer — and you've since moved into product, or you're ready to. You don't need a translator between you and the people building the system, and they don't need one between them and you.
What you'll do
- Own product direction across the full model lifecycle — data, training and adaptation, evaluation, release, production monitoring, and improvement or retirement — for our SLMs, ASR stack, and the real-time inference infrastructure that serves them. Retirement is a real part of that: the leading labs deliberately sunset models to concentrate effort, and we'd rather run a few models well than maintain a legacy model zoo. That's the whole job — not one rotation among many.
- Own the data strategy underneath it all: acquisition, consent and usage rights, sampling, and annotation. Model quality is decided here before the first training run — get the data model right and every ASR and SLM effort downstream gets simpler and better. On a platform built on customer conversations, consent and rights are foundational, not paperwork.
- Turn ambiguous model-quality questions into decisions. "Transcripts got worse this week" is a starting point, not a ticket. You'll define what good means, get it measured, and decide what ships.
- Sit inside eval reviews, error analyses, and incident retros as a peer. You should be able to look at a failing conversation trace and form your own hypothesis before the team tells you theirs.
- Treat internal teams as customers. The voice agents, agentic runtime, and AI features across the product all run on your stack — product engineering needs model capabilities and latency/cost envelopes they can plan around, and GTM needs a roadmap you won't have to walk back.
- Make trade-off calls with real constraints: model quality vs. streaming latency, train vs. fine-tune vs. buy, model size vs. capability, GPU cost vs. what the price point can absorb. These are the daily currency of this role, not edge cases.
- Write. Direction memos, decision docs, and specs that engineers actually read. If your best work happens in slide decks, this isn't the right fit.
- Release and rollback calls: whether a model ships, against quality bars you define.
- The model roadmap and its sequencing — including what gets deprecated and when.
- Where data investment goes: acquisition, annotation, and labeling priorities.
- The quality bar itself: what "good enough" means for an ASR or SLM release, and how it's measured.
And what you don't own, so there's no bait-and-switch: modeling and architecture choices belong to the engineers and researchers making them, and research bets and headcount are set with you, not by you. If a model regresses in production, accountability lands here first — the authority above is what makes that fair.
Skills you'll bring
- A hands-on track record with models themselves. You've built or run models in production — trained, fine-tuned, served, or optimized them — not just orchestrated APIs around them. Building beats running. Our core work is custom SLMs, ASR, and the inference infrastructure behind them, so experience with speech or with models under real-time constraints counts double. A CS/ML degree alone doesn't count; neither does "worked closely with data scientists."
- 2+ years of product ownership, formally titled or not. You've been accountable for what got built and whether it worked, not just for the backlog.
- Fluency across the model and serving stack. You have informed opinions on eval design, when to fine-tune vs. train vs. distill, quantization and serving trade-offs, why WER alone is a lousy ASR metric, and what actually drives real-time inference cost. Opinions you can defend to someone who does this full-time.
- Judgment under uncertainty. Model behavior is probabilistic; roadmaps aren't. You can commit to outcomes without pretending the uncertainty away.
- Direct communication. You say what you think, change your mind when the evidence says so, and put decisions in writing.
Nice to have
- Speech experience specifically: training or productionizing ASR/TTS, telephony, streaming latency work.
- You've built training data pipelines or run labeling operations — sourcing, sampling, annotation quality, data rights.
- You've run inference infrastructure at scale — GPU capacity planning, serving optimization, cost-per-call tuning.
- You've built or run an eval harness in production, not just read about them.
- Experience pricing or packaging AI products.
- Publications, open-source work, or a technical blog we can read before we talk.