Conversation, Person, Adult, Male, Man, Head, Computer Keyboard, Face, Coat, Monitor

Staff Software Engineer

 

Notice: Equinix is aware of scams involving fake employment offers. Read more. 

Staff Software Engineer

  • JR-162227
  • Hybrid
  • Bengaluru
  • Technology
  • Full time
View favorites

Who are we?

Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet. 
 

A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future.

Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.

Role Summary

We are seeking an Observability Engineer to manage, maintain, and support observability solutions across our cloud and application environments.


The role will focus on Grafana configuration and alerting, Infrastructure as Code using Terraform, and Python-based scripting and automation. You will work closely with Engineering, SRE, DevOps, and Platform teams to monitor systems, troubleshoot issues, and improve operational reliability.


Key Responsibilities

  • Manage and maintain Grafana dashboards, alerts, configurations, and monitoring views.

  • Configure and maintain monitoring and alerting based on application and infrastructure requirements.

  • Use Terraform for Infrastructure as Code (IaC) to manage and maintain observability-related infrastructure and configurations.

  • Develop and maintain Python scripts and automation to support monitoring, operational tasks, and reduce manual effort.

  • Monitor application and infrastructure health and identify gaps in observability and alerting.

  • Troubleshoot monitoring, alerting, and observability-related issues in production environments.

  • Support incident investigation and perform root cause analysis for recurring monitoring or infrastructure issues.

  • Work with Engineering, SRE, DevOps, and Platform teams to implement and maintain observability requirements.

  • Maintain documentation for dashboards, alerts, configurations, automation, and operational procedures.

  • Contribute to continuous improvements in monitoring, alerting, automation, and operational reliability.


Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.

  • 4–7 years of experience in Observability, SRE, DevOps, Infrastructure Engineering, Production Engineering, or a related field.

  • Hands-on experience with Grafana, including:

    • Dashboard creation and maintenance

    • Alert configuration and management

    • Grafana configuration and administration

  • Hands-on experience with Terraform and Infrastructure as Code (IaC).

  • Working experience with Python scripting and automation.

  • Good understanding of monitoring, alerting, metrics, and production support.

  • Working knowledge of Linux and cloud environments such as AWS or GCP.

  • Experience troubleshooting production environments and working with cross-functional engineering teams.


Preferred Qualifications

  • Experience with Prometheus and other monitoring platforms.

  • Exposure to Loki, Tempo, and OpenTelemetry.

  • Experience with Kubernetes and Docker.

  • Experience with ELK/OpenSearch or other logging platforms.

  • Experience with APM, distributed tracing, and application monitoring.

  • Experience with Java or Go.

  • Exposure to Helm or Ansible.

  • Experience integrating observability tools with PagerDuty, ServiceNow, Jira, Slack, or Microsoft Teams.

  • Experience working in a production support or on-call environment.


Key Skills

Mandatory: Grafana – Dashboards, Alerts & Configuration | Terraform – IaC | Python – Scripting & Automation

Observability: Monitoring, Alerting, Metrics, Logging, APM, Distributed Tracing

Cloud & Infrastructure: AWS/GCP, Linux, Kubernetes, Docker

Good to Have: Prometheus, Loki, Tempo, OpenTelemetry, ELK/OpenSearch, Java, Go, Helm, Ansible


What Success Looks Like

  • Grafana dashboards, alerts, and configurations are accurately maintained and reliable.

  • Observability and alerting requirements are implemented effectively across applications and infrastructure.

  • Terraform is used consistently to maintain observability-related infrastructure and configurations.

  • Python automation reduces repetitive operational work.

  • Monitoring and alerting gaps are identified and addressed proactively.

  • Production observability issues are investigated and resolved efficiently.

  • Engineering teams have reliable visibility into application and infrastructure health.

Equinix is committed to ensuring that our employment process is open to all individuals, including those with a disability.  If you are a qualified candidate and need assistance or an accommodation, please let us know by completing this form.

Equinix is an Equal Employment Opportunity and, in the U.S., an Affirmative Action employer.  All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law. 

We use artificial intelligence in our hiring process. Learn more here.

This posting is for a backfill position, meaning it is to fill an existing vacancy within our organization.