Director of Product Management, Agentforce Voice Models at Salesforce in San Francisco, CA
- Company: Salesforce
- Location: California - San Francisco
- Salary: $197K – $345K
- Job type: full time
- Workplace: onsite
- Posted: 2026-10-01
Job description
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Product Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce. About the Role Agentforce is Salesforce's next-generation AI platform, delivering autonomous agents that reason, take action, and communicate naturally across every customer touchpoint. Voice is the fastest-growing interaction surface—from contact-center automation and field-service assistants to real-time sales coaching and multilingual global deployments. As Director of Agentforce Voice Models, you will own the full voice intelligence stack: Automatic Speech Recognition (ASR / STT), Text-to-Speech (TTS), Speech-to-Speech (S2S) end-to-end pipelines, speaker diarization, prosody modeling, and the language-coverage roadmap that makes Agentforce sound natural in every market we serve. You will partner with product, infrastructure, and go-to-market teams to set the bar for transcription accuracy, latency, and voice expressiveness at enterprise scale. What You'll Do Voice Model Strategy & Roadmap Define and own the multi-year roadmap for Agentforce voice capabilities, spanning ASR/STT, TTS, S2S, and real-time voice agents. Set accuracy, latency, and quality benchmarks (WER, MOS, RTF, DMOS) and drive the organization to meet them. Evaluate build vs. buy vs. partner decisions for new voice model capabilities and maintain relationships with key academic and industry partners. ASR / STT (Automatic Speech Recognition) Lead the development of production-grade ASR systems optimized for telephony, WebRTC, and device-side deployment.Word Error Rate (WER) improvement across noise conditions, accents, and domain-specific vocabularyStreaming and batch recognition pipelines with sub-200 ms first-token latency targetsCustom vocabulary and language model adaptation (hot-word boosting, domain LM interpolation)Punctuation restoration and inverse text normalization (ITN) for downstream NLU Drive multilingual and code-switching ASR coverage across priority languages; govern the language onboarding process including data acquisition, model training, and acceptance testing. TTS (Text-to-Speech) Own the neural TTS pipeline—voice cloning, persona design, SSML compliance, and real-time synthesis—for Agentforce agent personas. Lead prosody research: intonation, rhythm, stress, and pause modeling that produces natural-sounding enterprise voices across conversational contexts. Manage voice talent agreements, ethical AI review, and consent frameworks for synthetic voice creation. Drive naturalness, expressiveness, and brand-consistency quality bars using subjective (MOS, CMOS) and objective (mel-cepstral distortion) evaluation frameworks. Speech-to-Speech (S2S) & Real-Time Voice Agents Architect low-latency S2S pipelines that enable full-duplex conversational AI without the ASR→NLU→TTS handoff penalty. Partner with the Agentforce Reasoning team to integrate voice understanding with agent action loops (tool calls, CRM lookups, escalation routing). Establish interruption, barge-in, and turn-taking models appropriate for enterprise voice agents. Speaker Diarization & Voice Analytics Deliver production speaker diarization ("who spoke when") for multi-party calls, enabling accurate per-speaker transcripts used in call coaching, compliance, and analytics. Develop speak
Director of Product Management, Agentforce Voice Models on JobPost.