HomeSearchCloud Engineer Jobs › Platform Engineer – Cloud & Observability

Platform Engineer – Cloud & Observability

Candescent

IN - Bengaluru - Office
More Cloud Engineer jobs: Cloud Engineer jobsCloud Engineer salary

Candescent is a forward-thinking technology company transforming how financial institutions deliver Intelligent Banking experiences. We unite digital banking, account opening, and branch solutions that power and connect digital banking, account opening, and branch solutions—creating seamless engagement across digital, remote, and in-person channels. Our Experience-Led, Intelligence-Driven approach combines human-centered design with data, automation, and cloud-based innovation.

Built on an API-first architecture, our extensible ecosystem enables institutions to adapt quickly, integrate easily, and unlock new opportunities for growth—turning every customer interaction into a moment of clarity, confidence, and connection. Position: Platform Engineer – Cloud(GCP) & Observability Location Bangalore Position Summary We are looking for a strong Platform Engineer  to help shape and evolve the enterprise platform that serves as the foundation for engineering teams across Candescent.

Leveraging GCP, Kubernetes, ArgoCD, Infrastructure as Code, automation, and modern observability practices , you will design and operate scalable, secure, reliable, and automated platform services that accelerate software delivery, improve reliability, and provide a consistent developer experience. The ideal candidate is passionate about platform ownership, observability, automation, reliability, and building systems that simplify operations at scale .

Key Responsibilities Define and implement observability standards  across applications and platform services, including metrics, logs, traces, dashboards, alerts, Golden Signals, SLIs, and SLOs. Design, build, and operate scalable, secure, and reliable platform services on GCP and Kubernetes/GKE . Develop reusable platform capabilities, standards, and self-service solutions that improve developer experience.

Implement and maintain Infrastructure as Code using Terraform  and automation using Python, Shell, or similar technologies. Support application delivery through ArgoCD, GitOps, and CI/CD  practices. Implement and integrate modern observability solutions for application and platform monitoring, logging, metrics, distributed tracing, alerting, and event correlation. Use observability data to troubleshoot complex production issues and identify reliability, performance, scalability, and capacity risks across cloud, platform, applications, networking, and service dependencies.

Participate in incident response, problem management, root cause analysis, and continuous improvement. Reduce operational toil through automation and self-service capabilities. Establish and improve platform standards, operational readiness, runbooks, and engineering practices. Collaborate closely with SRE, Software Engineering, Security, and Architecture teams to drive adoption of platform, observability, and reliability best practices.

Basic Requirements 6+ years of experience in Platform Engineering, Cloud Engineering, SRE, DevOps, Observability Engineering, or related disciplines . Strong hands-on experience with GCP and cloud architecture . Strong experience with Kubernetes/GKE  and cloud-native platforms. Hands-on experience with Terraform / Infrastructure as Code . Experience with ArgoCD, GitOps, and CI/CD . Strong understanding of cloud networking, IAM, compute, storage, and load balancing.

Strong understanding of observability concepts, including metrics, logs, distributed tracing, monitoring, alerting, Golden Signals, SLIs, and SLOs . Hands-on experience implementing or operating at least one modern observability platform such as Dynatrace, Datadog, Splunk, AppDynamics, Prometheus/Grafana, or similar . Experience designing dashboards, alerts, monitoring standards, and observability solutions for production environments.

Experience with automation using Python, Shell, or similar technologies. Strong production troubleshooting, incident management, and root cause analysis experience. Ability to troubleshoot issues holis

Search all live jobs — free, no account →