Bitdeer Group Logo

Bitdeer Group

AI Cloud Senior DevOps Engineer

Posted 3 Days Ago
In-Office
San Jose, CA
Senior level
In-Office
San Jose, CA
Senior level
Design, implement, and maintain CI/CD and MLOps pipelines, provision and scale cloud-native AI infrastructure (Kubernetes, GPU clusters), enforce IaC practices, implement observability and security frameworks, lead incident response and cross-functional platform engineering to ensure high availability and compliance.
The summary above was generated by AI

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/

Position Summary:

We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role, you will be the backbone of our deployment and infrastructure operations, ensuring that our AI-powered products and platforms are delivered with speed, security, and exceptional reliability. You will act as a crucial bridge between our research/development teams and real-world deployment, driving automation, optimizing cloud-native architectures, and establishing best practices for MLOps and traditional DevOps workflows.

Key Responsibilities:

  • CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
  • Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
  • High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
  • Infrastructure as Code (IaC): Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
  • Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics.
  • Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency through Internal Developer Platforms (IDP) and Platform Engineering initiatives.
  • Governance, Security & Compliance: Establish and enforce robust system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance with industry frameworks (e.g., SOC2, ISO27001).
  • Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, conduct thorough root cause analysis (RCA), and implement preventative remediation plans.

Basic Qualifications:

  • Experience & Education: Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
  • Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
  • Containerization & Orchestration: Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices.
  • Cloud Platforms: Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies.
  • Programming Skills: Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development.
  • Domain Knowledge: Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles.
  • Soft Skills: Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment.

Preferred Qualifications (Plus):

  • AI/ML Infrastructure Experience: Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads.
  • Large-Scale Systems: Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms).
  • Platform Engineering: Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity.
  • Advanced Security: Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001).
  • Leadership: Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Similar Jobs

21 Minutes Ago
Remote or Hybrid
220K-240K Annually
Mid level
220K-240K Annually
Mid level
Artificial Intelligence • Big Data • Software • Analytics • Business Intelligence • Big Data Analytics
This role involves driving revenue through new business development, managing sales cycles, and collaborating with various teams to meet goals.
Top Skills: AnalyticsDataDs/MlSaaSSQL
21 Minutes Ago
Hybrid
29-32 Hourly
Junior
29-32 Hourly
Junior
Artificial Intelligence • Big Data • Software • Analytics • Business Intelligence • Big Data Analytics
The Sales Development Representative will qualify leads, drive pipeline, and work closely with sales teams, focusing on mid-market and enterprise accounts.
Top Skills: Data AnalyticsSQL
21 Minutes Ago
Hybrid
178K-220K Annually
Senior level
178K-220K Annually
Senior level
Artificial Intelligence • Big Data • Software • Analytics • Business Intelligence • Big Data Analytics
The Product Manager will drive the vision for analytical features, work with customers, conduct market research, and collaborate with various teams to deliver impactful product features.
Top Skills: AIData Analytics

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account