Skip to main content

Protect Yourself from Recruitment Scams: Amgen is aware of fraudulent communications that may impersonate Amgen recruiters, employees, or recruiting partners. Official communications from Amgen will come from channels associated with the amgen.com domain. In some cases, Amgen may engage authorized third-party staffing or recruiting firms to contact candidates on our behalf. If you are contacted by a third-party recruiting firm about an Amgen opportunity, do not rely solely on the contact information provided in the message. Before sharing personal information, verify the opportunity by visiting the Amgen Careers website directly and contacting Amgen at AmgenCareers@careers.pure.cloud. You should also be informed in advance if an authorized recruiting firm will submit your resume to Amgen.

Amgen and its authorized recruiting partners will never request payment, gift cards, cryptocurrency, purchases, or banking information as part of an application or interview process. Be cautious of individuals falsely claiming to represent Amgen or an Amgen recruiting partner, and verify the authenticity of recruitment-related communications before responding.

Manager, Site Reliability Engineer - Data Platforms

Two lab technicians smiling
Buscar Empleos

SEE ALL JOBS

Manager, Site Reliability Engineer - Data Platforms

India - Hyderabad Apply Now
JOB ID: R-251352 País: India - Hyderabad Estado: On Site DATE POSTED: Sep. 24, 2026 CATEGORÍA DE EMPLEO: Engineering

Join Amgen’s Mission of Serving Patients

At Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do.

Since 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives.

Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career.

Manager, Site Reliability Engineer - Data Platforms

About Amgen

Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. 

About the Role

AWS Site Reliability Engineer with strong cloud architecture expertise to design, build, automate, and operate secure, resilient, scalable, and cost-efficient AWS platforms. You will combine AWS architecture leadership with practical Site Reliability Engineering-writing Infrastructure as Code, developing automation, improving observability, solving complex production problems, and strengthening operational excellence.

This role also includes responsibility for leading an SRE pod and managing assigned engineers. However, it is primarily a hands-on technical role. The successful candidate will remain a key technical contributor who leads through architecture, implementation, production ownership, incident response, coaching, and example.

What you will do

Roles & Responsibilities:

AWS Architecture and Platform Engineering

  • Design secure, highly available, scalable, and cost-efficient AWS architectures for EDSE applications and data platforms.
  • Define reusable platform services, Infrastructure as Code modules, reference architectures, and engineering guardrails.
  • Make and document architectural decisions across networking, identity, compute, containers, storage, security, observability, and resilience.
  • Lead architecture and production-readiness reviews, addressing reliability, security, performance, operability, and cost risks.

Reliability and Production Operations

  • Establish and improve SRE practices, including SLIs, SLOs, error budgets, availability targets, capacity planning, and operational-readiness criteria.
  • Use operational data to identify recurring failures, performance bottlenecks, capacity risks, and opportunities to reduce manual toil.
  • Participate in the on-call rotation and lead the technical response to complex or high-severity production incidents.
  • Conduct blameless post-incident reviews and ensure corrective actions produce durable engineering improvements.

Automation, Delivery, and Observability

  • Build reusable, tested infrastructure and operational automation using Terraform and programming languages such as Python.
  • Automate provisioning, configuration, validation, deployment, recovery, compliance checks, and routine operational activities.
  • Strengthen CI/CD and GitOps practices through automated testing, deployment controls, progressive delivery, and reliable rollback mechanisms.
  • Standardize metrics, logs, traces, dashboards, and actionable alerting to improve issue detection, diagnosis, and recovery.

Resilience, Security, and Cost Efficiency

  • Design and validate backup, high-availability, and disaster-recovery strategies aligned with defined business and recovery objectives.
  • Conduct restore tests, failover exercises, resilience reviews, and controlled game-day scenarios.
  • Embed least-privilege access, encryption, network segmentation, secrets management, vulnerability remediation, and policy-based controls into platform designs.
  • Improve AWS cost efficiency through right-sizing, resource lifecycle management, tagging, usage analysis, and architecture optimization without compromising reliability or security.

Pod and People Leadership

  • Lead the SRE pod by setting technical direction, prioritizing work, managing operational commitments, and driving delivery to completion.
  • Serve as a senior AWS and SRE advisor, partnering with application engineering, data engineering, cybersecurity, architecture, product, and FinOps teams.
  • Manage and mentor assigned engineers through regular feedback, one-on-one discussions, technical coaching, design and code reviews, and career-development support.

The expected allocation of responsibilities is:

  • Approximately 75-80% hands-on technical work, including AWS architecture, coding, Infrastructure as Code, automation, design reviews, production troubleshooting, observability, performance improvement, and incident response.
  • Approximately 20-25% pod and people leadership, including prioritization, work allocation, delivery coordination, mentoring, one-on-one discussions, performance feedback, career development, and removing team blockers.

The individual will be expected to move comfortably between architecture decisions, hands-on implementation, production operations, incident leadership, stakeholder communication, and people development. The balance may vary temporarily during major incidents, critical releases, or important delivery milestones.

What we expect of you

Basic Qualifications and Experience: 

  • Master’s or Bachelor’s degree in computer science or engineering field and 9 to 12 years of relevant experience, including substantial hands-on experience in AWS cloud engineering, platform engineering, infrastructure engineering, DevOps, or Site Reliability Engineering.
  • Prior experience leading a technical pod, engineering squad, or small team while continuing to contribute hands-on.

Must-Have Skills:

  • Significant hands-on experience designing, implementing, and operating business-critical production workloads on AWS, including experience with EKS, Sagemaker, Bedrock, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager,  Lambda, and RDS.
  • Deep knowledge of AWS architecture across networking, identity and access management, security, compute, containers, storage, monitoring, and resilience.
  • Strong experience with Infrastructure as Code using Terraform, CloudFormation, AWS CDK, or comparable technologies.
  • Proficiency in at least one programming or scripting language, such as Python, Go, Java, TypeScript, Bash, or PowerShell.
  • Ability to write maintainable, production-quality infrastructure code, automation, and operational tooling.
  • Practical experience applying SRE principles, including SLIs and SLOs, actionable alerting, incident response, root-cause analysis, post-incident improvement, and toil reduction.
  • Experience building or operating containerized platforms using Docker and Kubernetes, preferably Amazon EKS.
  • Experience with CI/CD, GitOps, source control, automated infrastructure testing, deployment controls, and safe production-change practices.
  • Experience implementing observability using AWS CloudWatch, OpenTelemetry, Prometheus, Grafana, Splunk, Datadog, or comparable platforms.
  • Strong knowledge of Linux, networking, distributed systems, performance analysis, high-availability design, and disaster recovery.
  • Demonstrated ability to troubleshoot complex production systems and lead technical decision-making during high-severity incidents.
  • Experience of FinOps practices, cloud cost allocation, unit-cost modeling, Savings Plans, and capacity optimization.
  • Experience leading an engineering pod, squad, or technical workstream and driving cross-functional delivery.
  • Experience managing or mentoring engineers, providing constructive feedback, and supporting their technical and career development.
  • Strong written and verbal communication skills, including the ability to explain technical risks and architecture trade-offs to engineering and non-engineering stakeholders.
  • Ability to collaborate effectively with global teams, make risk-based priority decisions, and drive complex work to completion.
  • Experience working in Agile delivery environments and using planning and delivery tools such as Jira or Jira Align.

Good-to-Have Skills:

  • Experience supporting enterprise data platforms or data-intensive workloads on AWS.
  • Experience with Databricks, Graph platforms, Lakehouse platforms, data lakes, or large-scale analytics environments. Candidates without prior Databricks experience should be willing to develop this capability after joining.
  • Experience with AWS Organizations, Control Tower, multi-account landing zones, service control policies, and cloud-governance frameworks.
  • Experience creating internal developer platforms, self-service infrastructure, golden paths, or paved-road capabilities.
  • Experience with policy as code, automated compliance, or security controls in regulated or highly governed environments.
  • Experience with chaos engineering, automated recovery, resilience testing, or large-scale performance testing.
  • Exposure to Azure, hybrid-cloud, or multi-cloud environments.
  • Relevant certifications are preferred but not required, including:
    • AWS Certified Solutions Architect - Professional
    • AWS Certified DevOps Engineer - Professional
    • SAFe Agilist or another SAFe certification

Functional Skills:

  • Excellent written and verbal communication, with the ability to explain complex platform concepts, architecture decisions, risks, and trade-offs in clear, business-relevant language.
  • Strong influencing and consensus-building skills, including the ability to establish standards and drive adoption across teams without relying solely on formal authority.
  • A platform-product mindset focused on reusable capabilities, paved roads, developer experience, measurable outcomes, and long-term platform health.
  • Strong systems-thinking and structured problem-solving skills, with the ability to diagnose issues across application, platform, cloud, governance, security, and operating-model boundaries.
  • High degree of ownership, initiative, and follow-through, with the ability to move ambiguous topics from exploration through decision, implementation, adoption, and continuous improvement.
  • Collaborative and globally minded, with experience working effectively across architecture, governance, cybersecurity, operations, engineering, product, and business teams.
  • Strong planning, estimation, prioritization, and execution skills, with the ability to manage multiple initiatives while maintaining high standards for security, reliability, quality, and reusability.
  • Ability to balance innovation with enterprise risk, distinguishing between experimentation, limited preview adoption, and production-ready capabilities.
  • Strong coaching and enablement skills, with the ability to create clear documentation, facilitate technical workshops, mentor engineers, and build an active platform community.
  • A growth mindset and commitment to continuous learning, modern engineering practices, constructive challenge, and responsible adoption of AI.
  • Increase the use of reusable, tested automation for infrastructure and operational changes.
  • Build a high-performing pod through clear priorities, strong engineering practices, effective coaching, and shared operational ownership.

What you can expect of us

As we work to develop treatments that take care of others, we also work to care for your professional and personal growth and well-being. From our competitive benefits to our collaborative culture, we’ll support your journey every step of the way.

In addition to the base salary, Amgen offers competitive and comprehensive Total Rewards Plans that are aligned with local industry standards.

Apply now

EQUAL OPPORTUNITY STATEMENT

Amgen is an Equal Opportunity employer and will consider all qualified applicants for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability status, or any other basis protected by applicable law.

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Ready to Apply for the Job?

We highly recommend utilizing Workday's robust Career Profile feature to complete the application process. A link to update your profile is available when you click Apply. You can then complete your Workday profile in minutes with the “Upload My Experience” functionality to upload an updated copy of your resume or you can simply edit the individual sections of your Career Profile.

Please note that you should be in your current position for at least 18 months before applying to internal positions. Staff must notify their current manager if invited for an interview. In addition, Staff are ineligible to apply for open positions if (a) their performance is currently being managed on a performance improvement plan (PIP) or other locally utilized formal coaching document or (b) their most recent performance rating was not a “Partially Meets Expectations” or higher. Please visit our Internal Transfer Guidelines for more detailed information

GCF Level

GCF Level 05

Career Category

Information Systems

Position Type

Full time

Apply Now
VIVE. GANA. PROSPERA.

Regístrate para recibir alertas de empleo

Mantente al día con las noticias y oportunidades de Amgen. Regístrate para recibir alertas sobre puestos que se adapten a tus habilidades e intereses profesionales.

Me interesa:Indique las primeras letras de una categoría y luego elija una a partir de las sugerencias. Después entre las primeras letras de un enlace y elija la opción que prefiera. Por último, haga clic en “Añadir” para crear su propia alerta.

  • Engineering, Hyderabad, State of Telangāna, IndiaBorrar

Al enviar tu información, reconoces que has leído nuestra política de privacidad (este contenido se abre en una nueva ventana) y consientes recibir comunicaciones por correo electrónico de.