Principal Software Engineer, Core Infrastructure
Oracle
Job at a glance
Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises worldwide. Our cloud offers hyperscale, multi-tenant services deployed across more than 50 regions globally. OCI continues to expand beyond traditional public-cloud boundaries to support dedicated, hybrid, and multicloud solutions, edge computing, and more. As a Principal Core Infrastructure Engineer, you will lead the design and evolution of foundational distributed systems behind OCI.
You will build highly scalable, elastic, and fault-tolerant services for high-volume data retrieval, storage, and processing, and set the technical direction for their reliability, correctness, security, and operational readiness. This position is office-based and requires onsite presence in Nashville, Tennessee. Relocation assistance may be available in accordance with Oracle's relocation policies.
What You'll Do Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale. Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them. Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design. Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts. Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.
Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness. Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence. Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management. Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.
What You'll Bring Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience. 7+ years of professional software-engineering experience, with demonstrated impact on large-scale distributed systems or cloud infrastructure. Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.
Proficiency in one or more object-oriented or systems programming languages, such as Java, C++, C#, or Go. Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices. Experience with system-level test automation, performance/load testing, reliability engineering, and production incident response. Demonstrated experience leading or influencing technical architecture and mentoring engineers.
Strong problem-solving, communication, and cross-functional collaboration skills. Preferred Qualifications Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms. Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing. Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.
Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure. Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range