Why This Job is Featured on The SaaS Jobs
This Multilingual AI Quality Specialist role sits at the intersection of product localisation and AI evaluation, a space that has become increasingly central as SaaS platforms ship generative and recommendation-driven experiences across markets. The remit signals a mature, platform-oriented environment where quality is treated as an operational capability, not a final review step, and where multilingual accuracy and cultural relevance directly influence user trust.
For a SaaS career, the work builds durable expertise in evaluation design, dataset curation, and quality signal creation, all of which translate across AI-powered products. It also develops the habit of connecting qualitative language judgments with measurable outcomes, strengthening the ability to influence roadmap decisions through evidence rather than opinion. Exposure to LLM evaluation patterns, including human-in-the-loop workflows, further aligns with how many SaaS companies are standardising AI release processes.
The role will suit professionals who enjoy structured methodology, documentation, and calibration, and who like partnering across Product, Engineering, Data Science, and market specialists. It fits someone comfortable operating in ambiguity while still imposing clear rubrics and thresholds, and someone motivated by improving systems at scale rather than owning a single locale or one-off QA cycle.
The section above is editorial commentary from The SaaS Jobs, provided to help SaaS professionals understand the role in a broader industry context.
Job Description
The Platform team creates the technology that enables Spotify to learn quickly and scale easily, enabling rapid growth in our users and our business around the globe. Spanning many disciplines, we work to make the business work; creating the infrastructure, tooling, frameworks, and capabilities needed to welcome a billion customers.
The Multilingual AI Data Quality space is part of Spotify's Global Language Quality Program within Localization. Bringing together multilingual language quality evaluation and AI data quality, the team helps ensure our AI-powered experiences are accurate, culturally relevant, and trustworthy across languages and markets.
Working closely with Product, Engineering, Data Science, Research, Personalization, Localization, vendors, and market experts, the team develops evaluation methodologies, high-quality datasets, and quality signals that help improve AI experiences and inform product decisions at scale.
\n
What You'll Do:
- Define quality frameworks, evaluation rubrics, thresholds, and methodologies for multilingual AI experiences.
- Design and execute structured evaluations for AI-generated, AI-translated, AI-curated, and recommendation-driven experiences.
- Lead multilingual dataset curation, annotation, enrichment, and ground-truth creation to support AI model development and evaluation.
- Analyze evaluation results, identify quality gaps, and provide actionable recommendations to improve multilingual AI quality.
- Support LLM-as-a-judge workflows, evaluator calibration, and human-AI agreement studies.
- Partner closely with Product, Engineering, Data Science, Research, Localization, vendors, and market experts to improve AI quality signals and inform launch decisions.
- Document best practices and help define quality standards across languages, markets, and AI use cases.
- Contribute to building scalable evaluation capabilities that support the next generation of AI-powered experiences across Spotify.
Who You Are:
- You have experience in multilingual quality evaluation, localization, data curation, annotation, AI evaluation, or related fields, including text-to-text and text-to-speech experiences.
- You understand language quality, cultural relevance, content quality, and user experience across multiple languages and markets.
- You have experience designing or conducting structured evaluations using quality rubrics, audits, annotation projects, or review methodologies.
- You are comfortable using qualitative and quantitative data to identify trends, measure quality, and make recommendations.
- You are familiar with large language models (LLMs), generative AI evaluation, human-in-the-loop workflows, or LLM-as-a-judge methodologies.
- You enjoy working through ambiguity and turning complex quality challenges into practical evaluation strategies.
- You communicate effectively and thrive in highly cross-functional environments, collaborating with technical and non-technical partners alike.
- Experience with recommendation systems, personalization, search, ranking, machine translation, generative AI, dataset creation, annotation operations, evaluator calibration, prompt testing, model evaluation, SQL, Python, dashboards, or annotation platforms is a plus.
Where You'll Be:
- This role is based in London or Stockholm.
- We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home.
\n
Spotify is an equal opportunity employer. You are welcome at Spotify for who you are, no matter where you come from, what you look like, or what’s playing in your headphones. Our platform is for everyone, and so is our workplace. The more voices we have represented and amplified in our business, the more we will all thrive, contribute, and be forward-thinking! So bring us your personal experience, your perspectives, and your background. It’s in our differences that we will find the power to keep revolutionizing the way the world listens.
At Spotify, we are passionate about inclusivity and making sure our entire recruitment process is accessible to everyone. We have ways to request reasonable accommodations during the interview process and help assist in what you need. If you need accommodations at any stage of the application or interview process, please let us know - we’re here to support you in any way we can.