Stellendetails
Revolutionierender Schutz.
Definieren Sie die Zukunft der Cybersicherheit.
Senior/Principal SRE Engineer
Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.Job Summary
Your Career:
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
- Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
- Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
- Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
- Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
- Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
- Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
- Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
- Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
- Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Qualifications
Your Experience:
- 5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
- Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
- Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
- Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
- Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
- Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
- Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
- Collaborative mindset - comfortable working across engineering, security, product, and leadership.
- Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
- Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
- Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
- Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
- Collaboration - You work effectively across engineering, security, product, and leadership to align on reliability priorities and drive shared outcomes.
- Innovation balanced with pragmatism - You actively explore new approaches, particularly AI-assisted operations and agent supervision, while weighing them against reliability, maintainability, and operational simplicity.
- Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
- Clear communication - You articulate reliability, risk, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
Our Commitment
We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.
We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.
All your information will be kept confidential according to EEO guidelines.
Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.Mehr zu Palo Alto Networks
-
Eine SaaS-Unternehmensgeschichte.
So hat Palo Alto Networks kritische SaaS-Apps mit SaaS Security Posture Management gesichert.
-
Unsere Kultur
Wegweisend in einer globalen Gemeinschaft – von der Vision zur Tat
-
Berufseinsteiger & Nachwuchsprogramme
Our early-in-career programs will train you to be a part of the next generation of cybersecurity talent.
Keine kürzlich angesehenen Jobs
Keine kürzlich angesehenen Jobs