Leads the strategy, modernization, reliability, automation, and performance of distributed VMware and Oracle Linux infrastructure platforms. Builds and mentors an SRE organization, drives Infrastructure as Code, observability, disaster recovery, high availability, incident management, and continuous improvement. Establishes SLOs, SLIs, error budgets, and operational standards while partnering with security, audit, application, network, storage, and executive stakeholders to deliver secure, scalable infrastructure.
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Director, Infrastructure & Site Reliability Engineering
Who is Mastercard?
Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential.
Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all.
Overview
Are you a visionary leader who thrives on driving transformation in complex infrastructure environments? Do you excel at building high-performing teams, fostering innovation, and aligning technology with business outcomes? The Distributed Platform Operations team is seeking a Director of Site Reliability Engineering (SRE) to lead strategic initiatives that ensure the reliability, scalability, and performance of our VMware and Oracle Linux platforms.
This role is ideal for a seasoned leader who combines deep technical expertise with a passion for operational excellence, automation, and cross-functional collaboration.
Key Responsibilities
• Define and execute the strategic roadmap for Site Reliability Engineering across distributed platforms.• Lead modernization efforts including hardware lifecycle management, virtualization upgrades, and infrastructure optimization.• Champion a culture of automation, resilience, and continuous improvement.• Build, mentor, and scale a high-impact SRE organization with a focus on technical excellence and career development.• Establish clear objectives, performance metrics, and development plans for team members.• Promote knowledge sharing and operational maturity through documentation and onboarding programs.• Oversee the health and performance of VMware clusters, ESXi hosts, and Oracle Linux environments.• Ensure robust disaster recovery and high availability strategies are in place and tested.• Drive incident management and root cause analysis for critical infrastructure issues.• Lead the adoption of Infrastructure-as-Code and automation frameworks using tools like Chef, Ansible, PowerCLI, Python, and Jenkins.• Reduce operational toil through scalable automation and self-healing systems.• Align engineering practices with DevOps principles and agile methodologies.• Architect observability solutions using Prometheus, Grafana, Splunk, and Dynatrace.• Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.• Optimize alerting and telemetry to support proactive incident response.• Ensure infrastructure compliance with security baselines, OS configurations, and regulatory standards.• Collaborate with InfoSec and audit teams to maintain a secure and compliant environment.• Partner with application, network, and storage teams to align infrastructure capabilities with business needs.• Communicate technical strategies, upgrade plans, and operational impacts to executive stakeholders.• Influence enterprise architecture and platform engineering decisions.
All about you
• 10+ years in Infrastructure, SRE, or Platform Engineering, with 5+ years in leadership roles • Strong expertise in VMware (ESXi, clusters) and Linux (preferably Oracle Linux) • Proven experience driving large-scale infrastructure modernization and automation initiatives • Hands-on experience with IaC and automation (e.g., Ansible, Chef, Python, Jenkins) • Solid understanding of SRE practices (SLOs, SLIs, error budgets, incident management) • Experience with observability tools (e.g., Prometheus, Grafana, Splunk, Dynatrace) • Strong knowledge of high availability, disaster recovery, and enterprise-scale infrastructure • Excellent leadership, stakeholder management, and executive communication skills
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Director, Infrastructure & Site Reliability Engineering
Who is Mastercard?
Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential.
Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all.
Overview
Are you a visionary leader who thrives on driving transformation in complex infrastructure environments? Do you excel at building high-performing teams, fostering innovation, and aligning technology with business outcomes? The Distributed Platform Operations team is seeking a Director of Site Reliability Engineering (SRE) to lead strategic initiatives that ensure the reliability, scalability, and performance of our VMware and Oracle Linux platforms.
This role is ideal for a seasoned leader who combines deep technical expertise with a passion for operational excellence, automation, and cross-functional collaboration.
Key Responsibilities
• Define and execute the strategic roadmap for Site Reliability Engineering across distributed platforms.• Lead modernization efforts including hardware lifecycle management, virtualization upgrades, and infrastructure optimization.• Champion a culture of automation, resilience, and continuous improvement.• Build, mentor, and scale a high-impact SRE organization with a focus on technical excellence and career development.• Establish clear objectives, performance metrics, and development plans for team members.• Promote knowledge sharing and operational maturity through documentation and onboarding programs.• Oversee the health and performance of VMware clusters, ESXi hosts, and Oracle Linux environments.• Ensure robust disaster recovery and high availability strategies are in place and tested.• Drive incident management and root cause analysis for critical infrastructure issues.• Lead the adoption of Infrastructure-as-Code and automation frameworks using tools like Chef, Ansible, PowerCLI, Python, and Jenkins.• Reduce operational toil through scalable automation and self-healing systems.• Align engineering practices with DevOps principles and agile methodologies.• Architect observability solutions using Prometheus, Grafana, Splunk, and Dynatrace.• Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.• Optimize alerting and telemetry to support proactive incident response.• Ensure infrastructure compliance with security baselines, OS configurations, and regulatory standards.• Collaborate with InfoSec and audit teams to maintain a secure and compliant environment.• Partner with application, network, and storage teams to align infrastructure capabilities with business needs.• Communicate technical strategies, upgrade plans, and operational impacts to executive stakeholders.• Influence enterprise architecture and platform engineering decisions.
All about you
• 10+ years in Infrastructure, SRE, or Platform Engineering, with 5+ years in leadership roles • Strong expertise in VMware (ESXi, clusters) and Linux (preferably Oracle Linux) • Proven experience driving large-scale infrastructure modernization and automation initiatives • Hands-on experience with IaC and automation (e.g., Ansible, Chef, Python, Jenkins) • Solid understanding of SRE practices (SLOs, SLIs, error budgets, incident management) • Experience with observability tools (e.g., Prometheus, Grafana, Splunk, Dynatrace) • Strong knowledge of high availability, disaster recovery, and enterprise-scale infrastructure • Excellent leadership, stakeholder management, and executive communication skills
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard's security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Similar Jobs at Mastercard
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Designs, develops, and configures ServiceNow solutions focused on IT asset management, software licensing, hardware management, and IT operations. Customizes workflows, business rules, and UI policies; integrates ServiceNow with external systems through REST and SOAP APIs; maintains platform performance and best practices; and collaborates with product owners and stakeholders to gather requirements and deliver solutions.
Top Skills:
Hardware Asset Management (Ham)HTMLIt Asset Management (Itam)It Operations Management (Itom)JAMFJavaScriptMicrosoft IntunePythonRest ApiSccmServicenowSoap ApiSoftware Asset Management (Sam)
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Owns the vision, roadmap, requirements, prioritization, release goals, and performance of technical platform features for Mastercard’s cross-border payments services. The role conducts customer and competitive research, translates business needs into feature definitions, makes data-driven trade-offs, partners with engineering and cross-functional teams, monitors delivery and operational health, and incorporates post-launch learnings into future planning.
Top Skills:
AgileAPIsCloud PlatformsSwift
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Drive business development and revenue growth for Mastercard Services across Commercial and New Payment Flows in North LAC. Manage existing client relationships, generate new opportunities, lead sales calls, qualify leads, develop solution proposals, manage pipelines, negotiate contracts, and achieve sales targets. Partner with internal teams on bids, account planning, client presentations, project delivery, and thought leadership while addressing complex client needs through data-driven, innovative solutions.
Top Skills:
Data AnalyticsInformation And Risk Management ServicesLoyalty Programs
What you need to know about the Los Angeles Tech Scene
Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

