Skip to content
Carrerlift

Sarvam

Machine Learning Engineer

Full-timePosted 1 month ago
Location
Bengaluru
Stipend
30-50 LPA
Type
Full-time

About this role

Batch: 2020/2021/2022/2023/2024/2025. What We're Looking For: - Strong Python and PyTorch — comfortable reading model internals, profiling inference, and debugging production failures - Hands-on experience integrating and optimising speech models (ASR or TTS) in production environments - Experience with real-time/streaming systems — WebSocket pipelines, chunked audio processing, or latency-sensitive async architectures - Solid understanding of modern speech system architectures — sequence-to-sequence models, attention mechanisms, flow-matching or diffusion-based TTS, streaming ASR - Familiarity with model serving infrastructure — Triton, TorchServe, ONNX Runtime, or equivalent - Experience with audio signal processing fundamentals: sample rates, PCM formats, spectrograms, vocoding, time-stretching - Strong async Python skills — asyncio, concurrent pipelines, managing backpressure in streaming systems - Comfort with ambiguity — the roadmap is not fully pre-specified - Undergraduate degree in a technical discipline (CS, EE, statistics, physics, or equivalent) Bonus Points: - Experience with multilingual or Indic speech systems — handling code-mixing, transliteration, tonal variation across Indian languages - Voice cloning or speaker adaptation techniques (zero-shot or few-shot) in production - Experience with real-time media protocols — WebRTC, RTMP, SRT, HLS, or real-time audio agent frameworks - Multi-participant audio systems — speaker diarization, concurrent pipeline management, per-user audio routing - Vocal source separation or speech enhancement techniques - Familiarity with FFmpeg, GStreamer, or media muxing/demuxing - Contributions to open-source speech/audio projects or a solid GitHub portfolio

How well do you fit this role?

Your résumé against this posting — what you have, what’s missing, and a short plan to close the gap.