Aalyria Logo

Aalyria

Staff Site Reliability Engineer - Spacetime

Posted 5 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
The Staff Site Reliability Engineer will design and manage Aalyria's centralized observability platform, focus on metrics, logging, and tracing systems, implement SLOs and SLIs, automate deployments, and drive incident response strategies for enhanced reliability across satellite and cloud platforms.
The summary above was generated by AI
About Aalyria:

Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. With technology acquired from Google, Aalyria is at the forefront of innovation in satellite and airborne mesh networks, as well as cislunar and deep-space communications. We are revolutionizing the orchestration and management of planetary mesh networks using any radio or optical spectrum, any orbit, and any hardware across land, sea, air, and space.

Role Overview:

This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space.

This is a greenfield/brownfield opportunity. You will be the foundational expert, defining the strategy and building the tools that empower our engineers. You will own the roadmap to mature our observability stack and build a robust, scalable, and insightful platform built on best-in-class technologies (e.g. Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to own a production-grade observability stack from the ground up, this role is for you.

Note: this role includes on-call responsibilities.

Key Responsibilities:
  • Design, build, and own the technical roadmap for Aalyria's centralized observability platform, integrating and scaling tools for metrics (Prometheus), logging (Loki), and distributed tracing (Tempo/OpenTelemetry).
  • Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
  • Establish and evangelize observability best practices, providing standards, documentation, and tooling (e.g., OpenTelemetry libraries) to empower our Go and Java application teams to instrument their services effectively.
  • Partner with core software engineers to provide the tools and insights needed to debug performance, optimize computational pipelines (including CPU/GPU workloads), and ensure the reliability of large-scale distributed systems.
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (Terraform) and GitOps principles (ArgoCD).
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
  • Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
Required Qualifications:
  • 7+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
  • Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
  • Proven mastery of Infrastructure as Code (IaC) with Terraform and GitOps principles (e.g., ArgoCD).
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services.
Preferred Qualifications:
  • Experience operating a multi-cloud environment, specifically GCP and AWS.
  • Hands-on experience with GitLab CI for CI/CD pipelines.
  • Working knowledge of service mesh technologies such as Istio or Linkerd.
  • Experience with high-performance computing (HPC) environments and instrumenting numerical optimization workloads
  • Familiarity with instrumenting applications written in Golang and C++.
  • Experience with JVM observability (tuning, monitoring) for Java-based applications.
What We Offer:
  • Innovative Environment: Work at a cutting-edge company shaping the future of aerospace communications.
  • Impactful Work: Directly contribute to critical national security programs and initiatives.
  • Growth Opportunities: Expand your career with opportunities for professional development and advancement.
  • Inclusive Culture: Be part of a collaborative, supportive, and inclusive workplace where your contributions matter.
  • Flexibility: Flexible working arrangements including hybrid remote/in-office schedules.
  • Compensation and Equity: Competitive salary, comprehensive benefits (401(k), dental, vision, health, life insurance), paid time off, and equity options.
ITAR/EAR Requirements:

This position involves access to export-controlled information. To comply with U.S. government export regulations, applicants must meet one of the following criteria:


(A) Qualify as a U.S. person, which includes:

  • U.S. citizen or national
  • U.S. lawful permanent resident (green card holder)
  • Refugee under 8 U.S.C. 1157
  • Asylee under 8 U.S.C. 1158

(B) Be eligible to access export-controlled information without requiring an export authorization.


(C) Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency.


The company reserves the right to decline pursuing an export licensing process for legitimate business-related reasons.

Equal Opportunity Employer Statement:

Aalyria is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex (including pregnancy, gender identity, and sexual orientation), national origin, age, disability status, genetic information, protected veteran status, or any other characteristic protected by law. Qualified applicants from all backgrounds are encouraged to apply.



Top Skills

AWS
Elk
GCP
Gitops
Go
Grafana
Jaeger
Java
Kubernetes
Loki
Opentelemetry
Prometheus
Python
Tempo
Terraform

Similar Jobs

An Hour Ago
In-Office or Remote
Canada, KS, USA
179K-211K Annually
Senior level
179K-211K Annually
Senior level
Blockchain • Fintech • Internet of Things • Cryptocurrency • Web3
As a Staff Site Reliability Engineer, you will lead the architectural design, operation, and scalability of cloud infrastructure while driving automation and ensuring reliability standards.
Top Skills: AWSAws CdkBashCrossplaneGoKubernetesPythonTerraform
8 Days Ago
Remote
United States
200K-250K Annually
Expert/Leader
200K-250K Annually
Expert/Leader
Fintech
As a Staff Site Reliability Engineer, you will shape reliability practices, optimize AWS infrastructure, lead incident response, and mentor engineers.
Top Skills: AWSDatadogGitopsTerraform
7 Days Ago
Remote
United States
Mid level
Mid level
Fitness
The Site Reliability Engineer will ensure system reliability and performance, design scalable architectures, improve CI/CD pipelines, maintain infrastructures, and lead incident response efforts.
Top Skills: ArgocdAWSDatadogDockerGithub ActionsGoJavaScriptKubernetesPrometheusPythonTerraform

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account