Oracle
Site Reliability Engineer
- Location
- PAN India
- Stipend
- Competitive
About this role
Batch: 2020/2021/2022/2023/2024. About the role: - Monitor, maintain and optimize the reliability of production systems and infrastructure - Respond to incidents and outages, diagnosing root causes and implementing fixes - Automate operational tasks to reduce manual work and improve efficiency - Collaborate with development and operations teams to improve system stability - Document runbooks and procedures for common operations and troubleshooting - Participate in on-call rotations to provide 24/7 support when needed What you'll work on: - Designing and implementing monitoring, logging and alerting systems - Managing deployment pipelines and ensuring reliable releases - Capacity planning and performance tuning of infrastructure - Building tools and scripts to streamline operational workflows - Conducting post-mortem analysis after incidents to prevent recurrence - Improving disaster recovery and business continuity processes What helps: - Strong programming skills in languages like Python, Go, Java or similar - Experience with Linux/Unix system administration - Knowledge of cloud platforms and containerization technologies - Familiarity with infrastructure-as-code and configuration management tools - Understanding of networking, databases and distributed systems - Problem-solving ability and attention to detail
How well do you fit this role?
Your résumé against this posting — what you have, what’s missing, and a short plan to close the gap.