Metasys Logo

Metasys

Site Reliability Engineer Internship

Reposted 6 Hours Ago
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
The summary above was generated by AI
Overview: Reliability and Operational Excellence

The Site Reliability Engineer (SRE) is responsible for the ultimate stability, performance, and scalability of our entire integrated supply chain e-commerce platform. You will apply software engineering principles to operations, ensuring the high availability and resilience of the customer-facing e-commerce storefront, internal SaaS tools (WMS, OMS), and specialized AI agent services.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will be the champion of uptime, performance, and automated operations for systems handling the critical MES → WMS → OMS flow.

  • Availability & SLO Management: Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for core business processes and all application layers. Manage the platform's overall Service Level Agreement (SLA).

  • Observability & Alerting: Architect, maintain, and optimize the comprehensive observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry). Develop high-fidelity alerting and ensure distributed tracing across the NestJS modular monolith and associated data stores (PostgreSQL, Redis).

  • Incident Response & Review: Own the incident response workflow, ensuring rapid triage, mitigation, and root cause analysis. Conduct thorough post-incident reviews to drive continuous improvement and eliminate recurring toil.

  • Scalability & Capacity Planning: Optimize auto-scaling policies for all services running on Docker containers. Conduct capacity planning based on business projections, especially for peak e-commerce and manufacturing load.

  • Disaster Recovery (DR): Design, implement, and regularly test Disaster Recovery procedures, including backup and restoration workflows for PostgreSQL 15 using tools like pgBackRest.

  • Automation: Eliminate operational toil through automation, managing infrastructure-as-code (Terraform) and CI/CD pipelines (Makefile).

Required Technologies & Tools

Candidates must possess deep experience in cloud operations, observability, and infrastructure automation:

  • Observability Stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (mandatory).

  • Infrastructure & Platform: Terraform, Docker, Traefik, Oracle Cloud Free VMs (or equivalent public cloud).

  • Data & Resilience: PostgreSQL (Deep knowledge), Redis, pgBackRest.

  • Automation: Strong scripting skills (Python/Bash) and experience with CI/CD tools and Makefile.

  • Methodology: Expert knowledge of SRE principles, toil reduction, and error budgeting.

AI Agent Focus

You will be responsible for the operational reliability of the emerging AI layer.

  • Agent Reliability: Implement specialized monitoring and logging for the AI agent services, ensuring LLM integrations and multi-agent systems (built with frameworks like LangChain) meet defined performance and availability SLOs.

  • Resource Optimization: Efficiently manage resource allocation for computationally intensive AI workloads to maintain platform stability and cost-efficiency.

Success Metrics & Career Path

Performance will be measured by:

  • Uptime/Availability: Achieving defined SLAs/SLOs across the platform.

  • MTTR: Reduction in Mean Time To Recover from production incidents.

  • Toil Reduction: Measured percentage reduction in manual, repetitive operational tasks through automation.

Mentorship Structure: Reports to the Head of Technology/CTO, working collaboratively with DevSecOps and development teams to ensure software is designed for reliability.

Similar Jobs

An Hour Ago
Remote or Hybrid
129K-233K Annually
Mid level
129K-233K Annually
Mid level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Own and grow Square’s Detroit territory through field-based prospecting, business visits, live product demonstrations, consultative selling, and full-cycle deal closing. Build pipeline through cold outreach, networking, events, referrals, and partnerships; develop relationships with local businesses; support onboarding; maintain Salesforce activity and forecasts; and consistently exceed sales quotas across Square’s software, hardware, and financial services products.
Top Skills: Payment Processing TechnologySalesforceSquare
An Hour Ago
Remote or Hybrid
CA, USA
164K-297K Annually
Senior level
164K-297K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Own the full outbound sales cycle for mid-market merchants, from prospecting and pipeline development through discovery, product demonstrations, negotiation, and close. Build net-new business through strategic outreach, sell multi-product solutions, manage complex multi-stakeholder deals, forecast accurately in Salesforce, and consistently exceed revenue targets. Collaborate with business development, product, marketing, implementation, and operations teams while serving as a consultative advisor to merchants.
Top Skills: Salesforce
An Hour Ago
Remote or Hybrid
CA, USA
95K-168K Annually
Entry level
95K-168K Annually
Entry level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Develop data science solutions for payments across Square and Cash App. Build AI-powered tools, ETL pipelines, dashboards, machine learning models, experiments, and cloud data systems. Partner with Product, Engineering, Finance, and Operations to improve payment success, cost efficiency, risk mitigation, and decision-making. Translate business needs into practical solutions, research questions independently, and operationalize data processes and infrastructure.
Top Skills: AICloud InfrastructureETLMachine LearningPythonSQL

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account