Conversation, Person, Adult, Male, Man, Head, Computer Keyboard, Face, Coat, Monitor

Senior Staff Engineer, Kubernetes Platform & Networking

 

Notice: Equinix is aware of scams involving fake employment offers. Read more. 

Senior Staff Engineer, Kubernetes Platform & Networking

  • JR-162794
  • Híbrido
  • Bengaluru
  • Technology
  • Full time
Ver favoritos

Who are we?

Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet. 
 

A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future.

Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.

Job Summary

We are looking for a Senior Staff Engineer to join the Reliability Engineering team that operates a bare-metal Kubernetes platform across multiple metro locations. The platform's architecture is defined and the build is underway, you will help deploy, operate and continuously improve it, from OS provisioning on racked servers through a multi-cluster fleet managed from a central management plane, with a software-defined data plane deployed via GitOps.

This is a deliberately broad role. We are not looking for someone who owns a layer and hands everything below it across a boundary. You will work alongside the teams that own the physical hardware, the underlay network, and the data plane software, and you need enough fluency in all three to reason about a fault before escalating it. Where you find the design falls short in practice, we expect you to say so and to shape how it evolves.

You will join the Digital Interconnection Engineering organization. The team owns the full application stack for interconnection and the Kubernetes platform for the software-defined data plane across multiple metros. We operate with an automation-first, GitOps-driven culture and a strong commitment to operational rigor, documented runbooks, and blameless postmortems.

Responsibilities

OS Provisioning & Bare-Metal Infrastructure

  • Own OS installation (Debian/Ubuntu) across bare-metal servers in multiple metros using PXE/iPXE imaging and cloud-init, and harden the OS prior to cluster bootstrap

  • Build and maintain a repeatable, zero-touch provisioning pipeline so every metro is deployed consistently without manual per-node intervention

  • Partner with the teams that own the hardware on BIOS/firmware baselines, high-speed NIC configuration and disk layout standards (separate NVMe for etcd, SSD for OS) and be able to diagnose faults at that layer, not just report them

Kubernetes Control Plane & Worker Nodes

  • Deploy and operate Kubernetes clusters across all metros with quorum-based control-plane fault tolerance and correct failure-domain placement across racks, PDUs, or availability zones

  • Maintain the Layer 4 load balancer or keepalived VIP fronting all kube-apiserver instances, so no client depends on a single node

  • Manage worker node provisioning, CNI configuration, and end-to-end cluster validation before workloads are onboarded

CI/CD & Data Plane Deployment

  • Build and own the CI/CD pipeline that deploys the software-defined data plane, including automated throughput and latency validation gates that must pass before promotion to production. The data plane software itself is configured and troubleshot by the team that owns it, your ownership is the pipeline and the gates

  • Own data-plane observability: monitoring, alerting thresholds, and operational runbooks

Network

  • Operate on top of a multi-metro underlay and partner with network and security teams on VLANs, BGP peering, NIC bonding, private connectivity, firewall rules and segmentation

  • Validate connectivity end to end across environments, and troubleshoot far enough into the network layer to isolate a problem before handing it over

OS & Kubernetes Upgrades

  • Plan and execute rolling OS and Kubernetes upgrades across the fleet with zero unplanned downtime; maintain tested upgrade runbooks and validated etcd snapshot/restore procedures before every major change

  • Own the etcd snapshot schedule, retention policy, and offsite storage; verify restore runbooks before each site goes live

  • Apply the platform's patching cadence across the fleet, tracking CVEs affecting the OS, Kubernetes and cluster components

Management Plane

  • Administer the centralized management plane across all clusters, including RBAC, cluster registration and policy enforcement, and keep production and non-production management planes strictly isolated to prevent blast-radius cascades from configuration changes

  • Drive GitOps-based fleet configuration management to keep clusters consistent, auditable, and drift-free

On-Call & Incident Management

  • Own your shift of a follow-the-sun on-call rotation, and lead incident response for control-plane, data-plane, and network events

  • Drive blameless postmortems and systematically eliminate recurring failure modes through automation, better runbooks, and sharper alerting

  • Contribute to the shared incident-command process, and maintain the runbooks for the systems you operate

What Success Looks Like, First 9-12 Months

  • Provisioning is fully automated and repeatable, new nodes and metros brought up with no manual per-node steps

  • Deployment pipeline in production, with throughput and latency gates enforced before promotion

  • Management plane administered across all clusters; production and non-production planes isolated; GitOps driving cluster configuration

  • OS and Kubernetes upgrades executed on cadence across the fleet with documented, tested runbooks

  • etcd snapshot and restore tested and signed off before each site goes live

  • Observability and alerting in place for the systems you operate, including the data plane

  • Recurring failure modes measurably reduced through automation and runbook improvements

Qualifications

Required

  • OS & bare metal: Deep hands-on Linux/Ubuntu provisioning at scale, PXE/iPXE imaging, cloud-init, BIOS/firmware management, OS hardening

  • Hardware fluency: Comfortable with BMC/IPMI, SMART data and platform sensors; able to diagnose disk, memory, NIC and power faults and work with vendors through warranty replacement

  • Kubernetes: Strong experience deploying and operating RKE2 or kubeadm clusters on bare metal, etcd operations, kube-apiserver HA, CNI (Multus/SR-IOV), multi-cluster fleet management

  • Rancher or equivalent management plane: Production experience managing multiple clusters through Rancher Prime, Rancher, or comparable fleet tooling

  • CI/CD for infrastructure: GitOps or pipeline-driven infrastructure deployment, Ansible, Terraform, ArgoCD, or Fleet

  • Networking: Hands-on with VLANs, BGP, bonded NICs, SR-IOV and private connectivity, with enough depth to troubleshoot across the boundary rather than only consume it

  • Upgrades: Demonstrated zero-downtime rolling upgrades of both OS and Kubernetes across a multi-node fleet

  • Operational rigor: Experience running platforms where availability is measured in minutes of downtime per year, and where an unreviewed change is the most likely cause of an outage

  • On-call: Comfortable owning an on-call shift and leading infrastructure incident response

  • Automation mindset: Everything-as-code, repeatable and auditable; strong bias toward automation over manual operations

Preferred

  • Familiarity with 6WIND VSR or other DPDK-based data planes, enough to build and validate a deployment and performance-testing pipeline around one

  • SR-IOV NIC tuning and data-plane performance validation

  • SDN or private-connectivity platforms (Equinix Fabric, VyOS, or equivalent)

  • Familiarity with LLM provider APIs and AI/agent gateway proxying patterns, since inference traffic from multiple providers transits this platform

  • Experience operating infrastructure across geographically distributed sites

Equinix is committed to ensuring that our employment process is open to all individuals, including those with a disability.  If you are a qualified candidate and need assistance or an accommodation, please let us know by completing this form.

Equinix is an Equal Employment Opportunity and, in the U.S., an Affirmative Action employer.  All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law. 

We use artificial intelligence in our hiring process. Learn more here.

This posting is a new position within our organization.