OXIO Logo

OXIO

Site Reliability Engineer

Posted 4 Days Ago
Remote
Hiring Remotely in USA
Entry level
Remote
Hiring Remotely in USA
Entry level
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
The summary above was generated by AI

Site Reliability Engineer
OXIO is the first NeoTelco. We are building the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network to millions of users.

 

We ensure that users and devices are connected, and stay connected wherever they go: Cross- country, carrier, or cellular technology. We help them pay less for mobile data. This technology is provided through our Carrier-as-a-Service platform: BrandVNO, a fully customizable telecom service. In addition, we enable clients of our service to extract the value from telecom data - enriching their customer experience, business intelligence, and product understanding in the many markets in which we operate.

 

Come join us in creating a modern technology platform with a group of engineers dedicated to advancing our vision. Our team is passionate about what we build, open to new ideas and challenges, and has our sights set on the future of connectivity.

 

Responsibilities

  • Design and implement platform on the cloud to support OXIO backend services

  • Automate technical operations: deployments, scaling, recovery, etc.

  • Monitor and maintain mission-critical production infrastructure to ensure maximum uptime

  • Participate in an on-call rotation and culture of continuous improvement through blameless postmortems

  • Enable the Engineering/Telecom/Data Engineering teams by providing them the tools to operate the service they build

 
Essentials
  • Understanding of Linux/Unix systems (most systems are Linux-based).

  • Familiarity with Linux/Unix system internals like process management, filesystems, memory management, and networking.

  • Proficiency in at least one programming language (Python, Go, or Ruby) and strong skills in scripting (Bash, Perl).

  • Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible.

  • Familiarity with containerization (Docker) and orchestration tools (Kubernetes).

  • Familiarity with monitoring tools like Prometheus, Grafana, or Datadog.

  • Knowledge of setting up alerts, analyzing logs, and creating dashboards for observability.

  • Familiarity with incident management practices (e.g., runbooks, postmortems).

  • Experience in being part of an on-call rotation and handling incidents.

  • Experience in setting up and maintaining Continuous Integration/Continuous Delivery pipelines (Jenkins, GitLab CI, CircleCI, etc.).

  • Hands-on experience with cloud providers (AWS, Google Cloud, Azure).

  • Knowledge of virtualization technologies (VMware, KVM) and cloud-native architecture.

  • Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls.

Nice to have
  • Strong understanding of deployment strategies (canary releases, blue-green deployments, etc.).

  • Familiarity with high availability and understanding failover mechanisms.

  • Familiarity with IAM (Identity and Access Management) and zero trust principles.

  • Experience working with distributed systems (e.g., Kafka, Cassandra, Elasticsearch).

  • Building custom monitoring tools or writing complex automation scripts.

  • Functional knowledge of database management (SQL and NoSQL).

  • Familiarity with distributed tracing (Jaeger, OpenTelemetry) and advanced log aggregation strategies (ELK stack, Splunk).

  • Familiarity with performance profiling tools and optimizing application performance under heavy load.

  • Familiarity in load testing and identifying bottlenecks.

  • Familiarity with Configuration Managment using SaltStack for maintaining server configurations.

Similar Jobs

3 Days Ago
Remote or Hybrid
OH, USA
Senior level
Senior level
Financial Services
Leads reliability engineering for mission-critical network services, including resiliency reviews, incident response, root-cause analysis, automation, observability, and durable remediation. Architects self-healing and guarded remediation workflows using Python, Shell, and Ansible. Provides technical leadership across SD-WAN, SDN, routing, switching, security, and traffic services while embedding SRE practices, resilience testing, AI-assisted workflows, and security controls. Mentors engineers and drives service-level objectives, error budgets, and operational readiness.
Top Skills: .NetAnsibleCi/CdContainer OrchestrationContainersFirewallsJavaLoad BalancersObservabilityProxiesPythonRoutingSd-WanSdaSdnShellSpring BootSwitching
4 Days Ago
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
10 Days Ago
In-Office or Remote
73K-130K Annually
Mid level
73K-130K Annually
Mid level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Architect, operate, and maintain resilient cloud infrastructure across Azure and AWS commercial and government environments. Support Kubernetes platforms, IaC, observability, monitoring, deployments, platform services, performance testing, and incident response. Define reliability metrics, participate in 24/7 on-call rotations, perform root cause analysis, and automate operational processes and remediation. The role requires U.S. citizenship and eligibility to obtain a Confidential, Secret, or Top Secret clearance.
Top Skills: ArgocdAWSAzureAzure MonitorDynatraceEncryptionFluxGitGitlabGitopsGrafanaHelmIaasIamKubernetesOwaspPaasPkiPrometheusPulumiRestful ApisSplunkTerraformVisual Studio Code

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account