Gallatin AI, Inc. Logo

Gallatin AI, Inc.

Machine Learning Operations (MLOps) Engineer

Posted 10 Days Ago
In-Office
El Segundo, CA, USA
80K-210K Annually
Senior level
In-Office
El Segundo, CA, USA
80K-210K Annually
Senior level
Build and operate production ML infrastructure across cloud, on-premises, GPU, edge, and disconnected environments. Own model training and serving, CI/CD, reproducibility, evaluation, observability, data pipelines, retrieval systems, and secure IL5/IL6 deployments. Support accreditation and classified environments while improving reliability, quality, and performance of LLM and ML systems.
The summary above was generated by AI
About Gallatin


At Gallatin, we are rebuilding logistics infrastructure for the national security missions of the United States and allied partners. We build AI systems that determine how logistics decisions are made — not just how they're executed. From factory to foxhole, we operate at the layer where data becomes decisions, and decisions make the advantage.


What You'll Do

In this role, you will build the systems that move a model out of a notebook and into the hands of a planner. Sometimes, that means deploying to an air-gapped rack in a tent instead of a VPC. You will own the path from training run to deployed capability: the infrastructure it runs on, the release process that ships it, the evaluation harness that proves it works, and the telemetry that tells us when it stops working.

Our AI/ML team works across retrieval-grounded systems for doctrine and logistics data, document and feature extraction, military symbol recognition, optimization and movement models, and an LLM agent platform. This role underpins that work: building the infrastructure, release processes, and evaluation systems that make it shippable and keep it honest in production. You will have plenty of room to shape how we build it.

Training & Serving Infrastructure
  • Own model training, fine-tuning, and batch inference infrastructure across AWS (SageMaker, EKS) and on-premises GPU hardware.

  • Stand up and tune LLM inference serving: vLLM-class stacks, quantization, continuous batching, KV-cache and throughput sizing. Make the hosted-vs-local call with numbers behind it.

  • Build for DDIL: local inference with configurable fallback, degraded-mode behavior, and sane resource envelopes on hardware we do not get to choose.

Release & Reproducibility
  • Build CI/CD for models and pipelines: versioned datasets, a model registry, promotion gates, and rollback that actually works under pressure.

  • Own infrastructure as code, containerization, and GitOps deployment across environments ranging from a dev cluster to a disconnected enclave.

  • Make reproducibility a hard requirement. Any result we put in front of a customer or evaluator must be reproducible from a commit and dataset version.

Evaluation & Observability
  • Build and own the evaluation harness: regression suites for retrieval and extraction pipelines, LLM-as-judge pipelines with measured judge-human agreement, and adversarial and held-out sets.

  • Instrument production for drift, latency, cost, retrieval quality, and failure modes, including quiet ones such as a retrieval miss that produces a fluent but wrong answer.

  • Make our metrics defensible to external test and evaluation reviewers. “We think it’s good” is not a deliverable.

Data & Pipeline Ownership
  • Own ingestion, versioning, and lineage for logistics and doctrinal data drawn from a heterogeneous set of authoritative sources.

  • Build and operate embedding and feature pipelines, incremental indexing, and the unglamorous systems that keep a retrieval index fresh.

  • Build the human-in-the-loop infrastructure: confidence-scored routing, review queues, and feedback capture that improves the next model.

Secure & Accredited Deployment
  • Deploy and operate ML systems in IL5 and IL6 environments, including air-gapped or restricted-network enclaves. Build the release, observability, artifact-management, and incident-response workflows those environments require.

  • Support ATO and continuous-authorization work with implementation evidence tied to applicable security controls (NIST SP 800-171, NIST SP 800-53 Rev. 5, CMMC Level 2, FIPS 140-3, and RMF/eMASS).

  • Handle CUI and classified data correctly without being asked twice.

What We’re Looking ForStrong Platform & Infrastructure Skills
  • 5+ years in MLOps, ML platform, or infrastructure engineering, including meaningful time working on systems with real users.

  • Strong Python skills and comfort in a production codebase, not just notebooks.

  • Deep Kubernetes and containerization experience, plus infrastructure as code.

  • Production experience with AWS ML/Azure infrastructure (SageMaker, EKS, or equivalent).

  • Hands-on GPU infrastructure experience: scheduling, utilization, memory sizing, and cost.

  • Hands-on experience deploying and operating production software in IL5 or IL6 environments, including disconnected or restricted-network deployments.

Production ML Judgment
  • You have shipped an LLM or ML system to production and then had to keep it working. You know what breaks.

  • You have built evaluation and monitoring for ML systems rather than adopting a vendor dashboard and hoping.

  • You can reason about where a pipeline’s quality actually comes from and say so when a metric is measuring the wrong thing.

Ownership
  • You are comfortable with ambiguity and owning a domain end to end. This is a small team; there is no one to hand the pager to.

  • You are willing to learn the mission domain. The engineers who do best here become genuinely interested in the logistics problem itself.

Nice to Have
  • Clearance: Preferred

  • LLM serving and inference optimization (vLLM, TensorRT-LLM, quantization, prefix caching).

  • Retrieval-grounded systems in production: hybrid retrieval, re-ranking, index freshness, and citation quality.

  • Edge, on-premises, or disconnected deployment.

  • Experience supporting ATO, continuous authorization, or production operations in classified environments.

  • Defense, intelligence, or another accredited or regulated environment.

  • Palantir Foundry, PostgreSQL/pgvector, NATS/JetStream, or ArgoCD.

  • A degree in CS, engineering, or a related technical field, or the equivalent built the hard way.


Mission and Identity

We are building the system that enables faster, smarter logistics decisions in contested environments, and we're doing it with a team of seasoned entrepreneurs, operators, and technologists who have built and scaled solutions in this space before. We hold ourselves to an extremely high standard. We value clear thinking, direct communication, and the kind of ownership that doesn't stop until something actually works.

Our mission is to create decision advantage when the stakes are the highest. If we succeed, the system doesn't just run; it gets smarter. We're not building AI for its own sake. We're building it because faster, smarter decisions in the most demanding environments on earth can't wait. If you want to work somewhere the stakes are real and the mission is urgent — you'll fit in here.


Why Gallatin?

The logistics infrastructure that supports America's warfighters and humanitarian disaster responders is overdue for transformation, and we are building it. From defense operations to disaster response, we're solving the hardest problems that keep missions moving when it matters most. Join a team where the mission is the point.
Compensation: Gallatin offers competitive compensation commensurate with experience. Actual compensation may vary based on experience, skills, and location. In addition to base salary, we offer a generous equity grant, full healthcare coverage, 401k, unlimited PTO, and the perks of working in a high-caliber, mission-driven environment.

Gallatin is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, national origin, age, disability, veteran status, sexual orientation, gender identity, or any other characteristic protected by applicable federal, state, or local law.

This position may require the ability to obtain and maintain a U.S. government security clearance. The successful candidate must be able to work in a classified environment when necessary.

We comply with the United States Department of Labor's Pay Transparency provision.

Note: Due to the nature of certain government contracts held by this organization, U.S. citizenship is a requirement for all positions at Gallatin. Proof of citizenship will be required prior to employment if selected.

Similar Jobs

9 Minutes Ago
In-Office
20-36 Hourly
Senior level
20-36 Hourly
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provides expert-level medical assistant care under provider supervision, including patient intake, venipuncture, EKGs, specimen collection, wound care, equipment sterilization, documentation, and infection control. Leads clinic improvement initiatives for 50% of the role, mentors and trains medical assistants, supports workflow coordination, maintains supplies, and promotes quality, confidentiality, and effective patient care.
Top Skills: AutoclaveCpt CodingDiagnostic EquipmentEkg EquipmentElectronic Health Records (Ehr)Icd-10Excel
9 Minutes Ago
In-Office
16-29 Hourly
Junior
16-29 Hourly
Junior
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provides non-clinical care management support in a high-volume remote healthcare environment. Responsibilities include managing member intake and admission or discharge information, processing referrals and prior authorizations, coordinating with providers and clinical teams, handling triage notifications, and resolving member or provider inquiries. Requires healthcare customer service experience, Cerner and medical terminology knowledge, Microsoft Office proficiency, and flexibility for four 10-hour shifts, rotating weekends, holidays, and occasional overtime.
Top Skills: CernerElectronic Medical RecordsExcelMicrosoft Word
9 Minutes Ago
In-Office
50K-89K Annually
Mid level
50K-89K Annually
Mid level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Supervise contact center employees by monitoring performance, coaching and conducting evaluations, managing staffing and schedules, hiring team members, resolving customer issues, monitoring service levels, addressing technical problems, supporting process improvements, and completing operational reports. The role requires flexible full-time availability, remote work, U.S. citizenship, and successful completion of a Department of Defense suitability or security clearance process.
Top Skills: Contact Center Systems And PlatformsIsoMicrosoft Office Suite

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account