Job Type: Full-Time Experience Required: 10 to 12 Years Location: Delhi ________________________________________ We are looking for an experienced and proactive NOC Manager/Leader to lead our 24x7 Network Operations Center team and ensure smooth, uninterrupted operations across our critical IT infrastructure. This role requires strong hands-on experience in Oracle Cloud Infrastructure (OCI), Linux, Kubernetes, Docker, cloud operations, monitoring & observability, security operations, incident management, automation, and automation frameworks.
The successful candidate will demonstrate strong leadership, technical troubleshooting capabilities combined with proven experience leading NOC/SOC teams, managing critical incidents, driving automation, improving operational reliability, and coordinating with cross-functional technology teams with a deep understanding of hybrid and multi-cloud environments. ________________________________________ Key Responsibilities: • Lead and manage a team of 24x7 NOC/SOC operations responsible for monitoring cloud and on-premises infrastructure, services, and applications.
• Ensure high availability, performance, and scalability across mission-critical systems. • Handle incident detection, response, resolution, and escalation, maintaining strict adherence to SLAs. • Execute and manage infrastructure and application deployments across OCI and hybrid environments. • Design and maintain alerting, logging, and monitoring solutions using Prometheus, Grafana, and OCI Monitoring.
• Automate repetitive operational tasks and infrastructure configurations using Ansible, Groovy, and Shell/Python scripting. • Maintain and troubleshoot Docker, Docker Swarm, and Kubernetes container environments. • Provide operational support and performance tuning for PostgreSQL and MySQL databases. • Maintain security, compliance, and patching of systems across environments. • Collaborate with cross-functional teams including DevOps, Development, and Support to drive continuous improvement in reliability and observability.
• Maintain accurate and timely incident documentation, and develop runbooks and SOPs for recurring events. • Provide leadership during critical incidents and participate in root cause analysis and preventive actions. • Participate in planning and implementation of DR and failover strategies. ________________________________________ Required Skills & Experience: • Strong Linux system administration knowledge.
• Proven experience in managing and operating Oracle Cloud Infrastructure (OCI). • Familiarity with other cloud platforms such as Azure or AWS is an advantage. • Solid experience with CI/CD tools such as Jenkins and related automation pipelines. • Hands-on expertise with Ansible for configuration management and automation. • Strong experience with Docker, Docker Swarm, and Kubernetes. • Proficient in monitoring and observability tools: Prometheus, Grafana, OCI Monitoring.
• Strong scripting skills in Shell, Python, and Groovy. • Operational experience managing PostgreSQL and MySQL databases in production. • Excellent communication and documentation skills. • Effective under pressure with strong problem-solving and decision-making capabilities. • Experience in a leadership or team lead role, managing 24x7 support environments. ________________________________________ Qualifications: • Bachelor’s degree in Computer Science, Information Technology, or related technical discipline.
• 5 to 7 years of experience in NOC, DevOps, or Cloud Operations roles. • Prior leadership experience is highly preferred. ________________________________________ Preferred Qualifications: • OCI Certification or equivalent certifications in AWS, Azure, or GCP. • Experience managing high-availability, large-scale, and mission-critical systems. • Experience managing NOC/SOC operations in a 24x7 environment.
• Experience implementing observability, automation, and self-healing operations. • Strong major incident management and execut