Flam
AI Engineer
- Location
- Bangalore
- Stipend
- 15-30 LPA
About this role
Batch: 2020/2021/2022/2023/2024. Required: - 2+ years building ML systems that ran in production, not only in notebooks - Strong Python and PyTorch (you can read a model implementation and modify it, not just call .fit) - Hands-on experience with at least one modern inference stack (vLLM, SGLang, TensorRTLLM, or TGI) and a real understanding of what makes it fast - Demonstrated fine-tuning experience — you've trained something, evaluated it properly, and know why your eval numbers meant what you claimed - Comfort with GPU-level reasoning: memory layout, precision trade-offs, where the bottleneck actually is - Ability to work from a paper (we move on recent research and expect you to be able to read it) Strongly Preferred: - Experience in one or more of: speech ASR/TTS, diffusion and image generation, or video/avatar generation - Quantization experience beyond running an off-the-shelf script - Real-time streaming systems (WebRTC, low-latency audio/video pipelines) - Indic language NLP, multilingual tokenizer work, or code-mixed data - Cloud GPU deployment on GCP, AWS, or RunPod — including the unglamorous parts
How well do you fit this role?
Your résumé against this posting — what you have, what’s missing, and a short plan to close the gap.