CoreWeave Logo

CoreWeave

Principal Engineer, Distributed Systems

Posted 44 Minutes Ago
Be an Early Applicant
In-Office
New York, NY
227K-303K Annually
Expert/Leader
In-Office
New York, NY
227K-303K Annually
Expert/Leader
Provide technical leadership for security-critical distributed systems, including authorization platforms, a Security Token Service, anomaly detection, bot defense, and API authentication gateways. Define multi-region architectures, fault-tolerance strategies, consistency models, security boundaries, observability, and operational standards. Guide architecture and design reviews, reliability and performance efforts, incident learning, and production readiness while mentoring senior engineers and aligning cross-functional teams.
The summary above was generated by AI
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.
About the role

We are looking for a Principal Engineer to provide technical leadership across Security Products. This is a senior individual-contributor role for an engineer who can define architecture, guide execution across multiple teams, and solve complex distributed-systems problems in security-critical infrastructure.

The systems you help design will support demanding production requirements: four nines of availability, high scalability, consistently low latency, strong security boundaries, and safe behavior under partial failure in a multi-region setup. You will work across the full lifecycle of these systems, from architecture and technical strategy through implementation guidance, operational readiness, incident learning, and long-term evolution.

Your work will span across high-scale authorization systems, anomaly detection systems, bot-defense and security token service (STS), API authentication gateway, and the shared infrastructure required to operate these capabilities reliably across regions and deployment environments.

You will partner with engineering, security, infrastructure, networking, platform, product, and customer-facing teams. Success in this role requires both deep technical judgment and the ability to create alignment, raise engineering standards, and make complex architecture understandable and actionable for others.

What you will do
  • Establish architectural approaches for distributed systems - multi-region operation, including service placement, failover, replication, traffic management, disaster recovery, and regional independence.
  • Drive decisions around consistency models, caching, invalidation, propagation, revocation, idempotency, concurrency, and ordering where correctness and security are critical. 
  • Design for data residency, tenant isolation, trust boundaries, blast-radius reduction, and controlled handling of sensitive security data.
  • Define fault-tolerance strategies for dependencies, networks, regions, storage systems, and control-plane components, including graceful degradation and safe recovery.
  • Design systems that meet four-nines availability goals while maintaining predictable low latency and high throughput under normal operation, traffic spikes, and partial failures.
  • Improve the reliability and operability of critical services through SLOs, error budgets, metrics, logs, traces, audit events, alerting, incident response, and post-incident learning.
  • Guide teams through architecture reviews, design reviews, implementation tradeoffs, capacity planning, load testing, performance analysis, and production readiness assessments.
  • Mentor senior and staff engineers, develop technical talent, and raise the quality of engineering practice across the organization.
  • Communicate architecture, tradeoffs, risks, and recommendations clearly to technical and executive audiences.
  • Lead the design of high-scale authorization systems that support policy authoring, policy evaluation, access-control enforcement, auditability, and integration across many services and tenants.
  • Provide technical direction for a Security Token Service, including token issuance, validation, lifecycle management, trust relationships, key rotation, revocation, and secure service-to-service access.
  • Guide the design and implementation of an API Authentication Gateway that provides consistent, secure, observable authentication and authorization for CoreWeave APIs.
  • Set technical standards for API design, service contracts, threat modeling, cryptographic key management, secrets handling, observability, and operational readiness.
  • Partner with service teams to make authorization and authentication capabilities easy to adopt through clear interfaces, SDKs, reference implementations, documentation, and reliable integration patterns.
What you bring
  • 12+ years of experience designing and building production software, including substantial experience with distributed systems and platform infrastructure.
  • A track record of serving as a principal, staff-plus, distinguished, or equivalent technical leader across multiple teams or a broad technical domain.
  • Deep expertise in distributed-systems design, including scalability, availability, latency, consistency, concurrency, partition tolerance, caching, replication, and failure recovery.
  • Experience designing and operating security-sensitive or reliability-critical services in production.
  • Strong software engineering experience with one or more systems languages or backend ecosystems, such as Go, Java, Rust, C++, Python, or equivalent.
  • Demonstrated ability to move from ambiguous requirements to clear architecture, sequenced execution plans, and durable engineering outcomes.
  • Experience making architecture decisions that balance security, correctness, performance, operational simplicity, and time to value.
  • Strong written and verbal communication skills, including the ability to influence without direct authority and build alignment across organizational boundaries.
  • A track record of improving engineering standards, mentoring senior engineers, and helping teams deliver complex systems safely.
Preferred qualifications
  • Experience designing and implementing authorization systems at high scale, including policy engines, permission models, policy distribution, decision services, enforcement points, or authorization observability.
  • Experience designing or operating a Security Token Service, identity platform, credential service, or comparable security-critical control plane.
  • Experience designing or operating an API Authentication Gateway, service-mesh authorization layer, or centralized authentication and authorization platform.
  • Experience operating systems with four-nines availability requirements, high request volume, strict latency objectives, and demanding reliability expectations.
  • Experience with multi-region architecture, regional failover, active-active or active-passive operation, replication, disaster recovery, and regional isolation.
  • Experience reasoning about strong, eventual, and bounded-staleness consistency models, especially for security policy, credentials, tokens, permissions, and revocation.
  • Experience designing for data residency, regional data controls, tenant isolation, trust-domain separation, and compliance-sensitive deployment models.
  • Experience building systems that remain secure and useful during dependency failures, network partitions, regional outages, stale data, degraded capacity, or partial compromise.
  • Experience with authentication and identity standards such as OAuth 2.0, OpenID Connect, SAML, SCIM, JWT, mTLS, SPIFFE/SPIRE, or related protocols.
  • Experience with cryptography, key management, HSMs, secrets management, certificate authorities, signing keys, token validation, or secure credential lifecycle management.
  • Experience with Kubernetes, cloud infrastructure, service networking, distributed storage, messaging systems, and multi-cluster operations.
  • Experience building security and platform services with strong observability, including metrics, logs, traces, audit records, security events, and actionable alerting.
  • Experience working with enterprise customers, regulated workloads, or hybrid and multi-cloud environments.
How you will be successful

In your first year, you will:

  • Establish a clear technical north star for Security Products’ authorization, authentication, and security-infrastructure capabilities.
  • Build strong working relationships with engineering and product leaders, staff-plus engineers, security teams, and the service teams that depend on these platforms.
  • Advance the architecture and delivery plan for high-scale authorization systems, the Security Token Service, and the API Authentication Gateway.
  • Improve the reliability, performance, observability, and operational maturity of critical security services against four-nines availability goals.
  • Help teams make explicit, durable decisions about multi-region design, data consistency, data residency, tenant isolation, fault tolerance, and security boundaries.
  • Reduce duplicated access-control and authentication patterns by enabling consistent, well-documented platform capabilities.
  • Raise the quality of architecture reviews, threat modeling, performance testing, capacity planning, and production readiness across Security Products.
  • Mentor senior engineers and help grow a strong technical leadership community.
Why join Security Products

Security Products is building the systems that make CoreWeave trustworthy at scale. You will work on foundational infrastructure that protects access to cloud resources, enables secure product experiences, and supports customers running critical AI workloads.

This is an opportunity to shape the architecture of a modern cloud security platform while solving problems at the intersection of distributed systems, security engineering, reliability, performance, and global-scale infrastructure.

Location requirement

This position is based in Sunnyvale, California or New York, New York and requires working on-site. Remote work is not available for this role.

CoreWeave is an equal opportunity employer. We evaluate qualified applicants without regard to legally protected characteristics.

The base salary range for this role is $227,000 to $303,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). 

What We Offer

The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance 
  • Voluntary supplemental life insurance 
  • Short and long-term disability insurance 
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement 
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health 
  • Family-Forming support provided by Carrot
  • Paid Parental Leave 
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

California Applicants

California Consumer Privacy Act 

Equal Opportunity & Accommodations

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: [email protected].

Export Control Compliance

This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency.  CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

Similar Jobs at CoreWeave

2 Hours Ago
In-Office
157K-210K Annually
Senior level
157K-210K Annually
Senior level
Cloud • Information Technology • Machine Learning
Own infrastructure capacity planning for hardware New Product Introduction programs. Track power, cooling, network, and space roadmaps; determine hardware ordering cutovers; communicate scope and BOM requirements; and drive cross-functional readiness across Capacity, Supply Chain, Construction, Design, and Engineering. Translate new GPU, CPU, liquid-cooling, and rack hardware requirements into scalable data center designs, manage milestones and risks, support lifecycle transitions, and ensure infrastructure is ready for deployment.
Top Skills: Coolant Distribution Units (Cdus)Cooling InfrastructureCpusData Center InfrastructureDcim PlatformsGpusLiquid CoolingNetwork InfrastructurePower Infrastructure
2 Hours Ago
In-Office
143K-210K Annually
Senior level
143K-210K Annually
Senior level
Cloud • Information Technology • Machine Learning
Leads technical due diligence for prospective data center sites, evaluating electrical, mechanical, civil, structural, architectural, utility, zoning, and infrastructure feasibility. Reviews provider proposals, capacity data, conceptual layouts, and building documentation; identifies risks, upgrades, cost and schedule impacts; prepares go/no-go recommendations and technical assessments. Coordinates with utilities, landlords, developers, internal subject matter experts, and design teams, then hands secured sites to Design Managers. Frequent travel is required.
Top Skills: CdusChilled Water PlantsDashboardsIt/LvMepPower DistributionSubstationsTcs LoopsUtility Infrastructure
2 Hours Ago
In-Office
188K-275K Annually
Senior level
188K-275K Annually
Senior level
Cloud • Information Technology • Machine Learning
Own the strategy, roadmap, and execution for W&B Launch, infrastructure insights, and infrastructure-native developer tools integrating W&B with CoreWeave. Partner with frontier AI teams, platform engineering, and W&B engineering to improve job submission, compute orchestration, observability, and automated research workflows. Drive product concepts from discovery through launch, translate infrastructure needs into developer experiences, and measure and iterate on product impact.
Top Skills: Api DesignDistributed TrainingExperiment TrackingGpu ComputingHyperparameter OptimizationJaxKubernetesLlm DevelopmentModel ManagementPyTorchSlurmW&B Launch

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account