Senior Staff Engineer, Kubernetes Platform & Networking
Notice: Equinix is aware of scams involving fake employment offers. Read more.
Senior Staff Engineer, Kubernetes Platform & Networking
- JR-162794
- Hybride
- Bengaluru
- Technology
- Full time
Who are we?
Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet.
A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future.
Job Summary
We are looking for a Senior Staff Engineer to join the Reliability Engineering team that operates a bare-metal Kubernetes platform across multiple metro locations. The platform's architecture is defined and the build is underway, you will help deploy, operate and continuously improve it, from OS provisioning on racked servers through a multi-cluster fleet managed from a central management plane, with a software-defined data plane deployed via GitOps.
This is a deliberately broad role. We are not looking for someone who owns a layer and hands everything below it across a boundary. You will work alongside the teams that own the physical hardware, the underlay network, and the data plane software, and you need enough fluency in all three to reason about a fault before escalating it. Where you find the design falls short in practice, we expect you to say so and to shape how it evolves.
You will join the Digital Interconnection Engineering organization. The team owns the full application stack for interconnection and the Kubernetes platform for the software-defined data plane across multiple metros. We operate with an automation-first, GitOps-driven culture and a strong commitment to operational rigor, documented runbooks, and blameless postmortems.
Responsibilities
OS Provisioning & Bare-Metal Infrastructure
Own OS installation (Debian/Ubuntu) across bare-metal servers in multiple metros using PXE/iPXE imaging and cloud-init, and harden the OS prior to cluster bootstrap
Build and maintain a repeatable, zero-touch provisioning pipeline so every metro is deployed consistently without manual per-node intervention
Partner with the teams that own the hardware on BIOS/firmware baselines, high-speed NIC configuration and disk layout standards (separate NVMe for etcd, SSD for OS) and be able to diagnose faults at that layer, not just report them
Kubernetes Control Plane & Worker Nodes
Deploy and operate Kubernetes clusters across all metros with quorum-based control-plane fault tolerance and correct failure-domain placement across racks, PDUs, or availability zones
Maintain the Layer 4 load balancer or keepalived VIP fronting all kube-apiserver instances, so no client depends on a single node
Manage worker node provisioning, CNI configuration, and end-to-end cluster validation before workloads are onboarded
CI/CD & Data Plane Deployment
Build and own the CI/CD pipeline that deploys the software-defined data plane, including automated throughput and latency validation gates that must pass before promotion to production. The data plane software itself is configured and troubleshot by the team that owns it, your ownership is the pipeline and the gates
Own data-plane observability: monitoring, alerting thresholds, and operational runbooks
Network
Operate on top of a multi-metro underlay and partner with network and security teams on VLANs, BGP peering, NIC bonding, private connectivity, firewall rules and segmentation
Validate connectivity end to end across environments, and troubleshoot far enough into the network layer to isolate a problem before handing it over
OS & Kubernetes Upgrades
Plan and execute rolling OS and Kubernetes upgrades across the fleet with zero unplanned downtime; maintain tested upgrade runbooks and validated etcd snapshot/restore procedures before every major change
Own the etcd snapshot schedule, retention policy, and offsite storage; verify restore runbooks before each site goes live
Apply the platform's patching cadence across the fleet, tracking CVEs affecting the OS, Kubernetes and cluster components
Management Plane
Administer the centralized management plane across all clusters, including RBAC, cluster registration and policy enforcement, and keep production and non-production management planes strictly isolated to prevent blast-radius cascades from configuration changes
Drive GitOps-based fleet configuration management to keep clusters consistent, auditable, and drift-free
On-Call & Incident Management
Own your shift of a follow-the-sun on-call rotation, and lead incident response for control-plane, data-plane, and network events
Drive blameless postmortems and systematically eliminate recurring failure modes through automation, better runbooks, and sharper alerting
Contribute to the shared incident-command process, and maintain the runbooks for the systems you operate
What Success Looks Like, First 9-12 Months
Provisioning is fully automated and repeatable, new nodes and metros brought up with no manual per-node steps
Deployment pipeline in production, with throughput and latency gates enforced before promotion
Management plane administered across all clusters; production and non-production planes isolated; GitOps driving cluster configuration
OS and Kubernetes upgrades executed on cadence across the fleet with documented, tested runbooks
etcd snapshot and restore tested and signed off before each site goes live
Observability and alerting in place for the systems you operate, including the data plane
Recurring failure modes measurably reduced through automation and runbook improvements
Qualifications
Required
OS & bare metal: Deep hands-on Linux/Ubuntu provisioning at scale, PXE/iPXE imaging, cloud-init, BIOS/firmware management, OS hardening
Hardware fluency: Comfortable with BMC/IPMI, SMART data and platform sensors; able to diagnose disk, memory, NIC and power faults and work with vendors through warranty replacement
Kubernetes: Strong experience deploying and operating RKE2 or kubeadm clusters on bare metal, etcd operations, kube-apiserver HA, CNI (Multus/SR-IOV), multi-cluster fleet management
Rancher or equivalent management plane: Production experience managing multiple clusters through Rancher Prime, Rancher, or comparable fleet tooling
CI/CD for infrastructure: GitOps or pipeline-driven infrastructure deployment, Ansible, Terraform, ArgoCD, or Fleet
Networking: Hands-on with VLANs, BGP, bonded NICs, SR-IOV and private connectivity, with enough depth to troubleshoot across the boundary rather than only consume it
Upgrades: Demonstrated zero-downtime rolling upgrades of both OS and Kubernetes across a multi-node fleet
Operational rigor: Experience running platforms where availability is measured in minutes of downtime per year, and where an unreviewed change is the most likely cause of an outage
On-call: Comfortable owning an on-call shift and leading infrastructure incident response
Automation mindset: Everything-as-code, repeatable and auditable; strong bias toward automation over manual operations
Preferred
Familiarity with 6WIND VSR or other DPDK-based data planes, enough to build and validate a deployment and performance-testing pipeline around one
SR-IOV NIC tuning and data-plane performance validation
SDN or private-connectivity platforms (Equinix Fabric, VyOS, or equivalent)
Familiarity with LLM provider APIs and AI/agent gateway proxying patterns, since inference traffic from multiple providers transits this platform
Experience operating infrastructure across geographically distributed sites
Equinix is committed to ensuring that our employment process is open to all individuals, including those with a disability. If you are a qualified candidate and need assistance or an accommodation, please let us know by completing this form.
Equinix is an Equal Employment Opportunity and, in the U.S., an Affirmative Action employer. All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law.
We use artificial intelligence in our hiring process. Learn more here.
This posting is a new position within our organization.