Startup Talents Logo

Startup Talents

Staff Software Engineer

Posted 8 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Expert/Leader
Remote
Hiring Remotely in USA
Expert/Leader
Own backend and AI infrastructure for autonomous trading agents. Build the Go-based agent runtime, execution and data layers, real-time trading systems, deployment platform, model hosting migration, telemetry, CI/CD, monitoring, and incident response. Develop production services primarily in Go, Python, and TypeScript while operating AWS EKS infrastructure. The role requires end-to-end ownership of reliable, low-latency systems managing real-time trading and financially sensitive positions.
The summary above was generated by AI

This is a remote position.

The company is an OnchainGPT for autonomous trading. It is an AI-powered mentor, advisor, and companion that helps users explore the onchain world, find alpha, auto-trade, and outsmart the market. 

They are looking for Staff Software Engineer (Backend and AI Infra) who will own two critical workstreams :
- the agent runtime and backend infrastructure that powers every trade in their fleet,
- and the migration of model hosting and agent deployment in-house, moving the company off third-party LLM providers and hosted agent platforms to own infrastructure. 


NOTE : This role is NOT a pure Devops role but rather a building role in which you will spend 80% of your time writing Go, Python, and TypeScript that ships to production. 

You’re writing the backend services, runtime engine, and deployment systems that the entire company's agent fleet runs on.
When you ship, every agent in the fleet immediately gets faster, more reliable, and more autonomous.
You are in charge of managing the very infrastructure you are building, as its designer you are be the best person to operate it.
You’re building the backend for autonomous AI agents that manage real money in real time.
The runtime you build determines whether positions are protected.
The model hosting you stand up determines whether agents can think.
The deployment pipeline you create determines whether the fleet can evolve.
This is foundational infrastructure for a new category of software.


MISSIONS





Agent Runtime & Backend (+/- 50%)

The runtime is the engine that makes every agent work. You’ll own the core systems:

Plugin Runtime — the per-agent process that runs position tracking (10s polling), the RatchetStop exit engine (tiered trailing stops with sub-second evaluation), and DSL state management. Currently Go + Python; migrating to a centralized Go service with Postgres state and real-time websocket price feeds

Scanner Gateway / Rules Engine — a YAML-configurable evaluation layer that sits between scanners and execution. Scanners produce raw signal variables; the rules engine applies gates, scoring, and filters defined in YAML. Users customize trading behavior without touching Python. This is the next major runtime feature

RatchetStop Backend — centralized profit-trailing service that protects positions even when the agent is offline. Evaluates tier upgrades and places stop-loss orders on Hyperliquid via websocket, replacing per-agent polling with condition-based evaluation across all positions

Execution Layer — the MCP (Model Context Protocol) server that bridges agents to the company's 48+  platform tools: position creation, clearinghouse state, market data, Smart Money intelligence. You’ll own auth, rate limiting, and the contract between agents and the exchange

Data Layer — enriched Hyperfeed pipeline (top 1K trader positions, momentum events, market concentration) flowing through Redis, Postgres, and ClickHouse. Real-time ingestion, 4-hour rolling windows, and the APIs that every scanner calls

Model & Agent Hosting Migration (+/- 30%)

The company is moving off third-party hosted agents and external LLM inference to the company's owned infrastructure. You’ll lead the technical execution:

Agent deployment platform — migrate agents from Railway/OpenClaw to own-hosted infrastructure. Each agent needs isolated workspace, cron scheduling, state persistence, MCP connectivity, and Telegram notifications. Target: deploy any skill from a GitHub repo with one command

Model hosting — evaluate and implement the path from external LLM APIs (Anthropic, Google) to self-hosted inference. Options range from proxied external models with full telemetry capture, to fine-tuned models running on own GPUs. You’ll own the decision and execution

Agent telemetry — capture every scanner evaluation, every trade decision, every signal score across all agents. This data feeds the self-reinforcing loop: agents learn from fleet-wide performance, fork winning strategies, and improve autonomously

Deployment pipeline — CI/CD for shipping scanner updates, runtime patches, and skill configs to 50+ live agents without interrupting open positions. Zero-downtime rollouts where downtime = unprotected capital

Infrastructure & Operations (+/- 20%)

Build monitoring and alerting that catches agent failures, orphaned positions, state corruption, and Auth expiration before they cost money

Manage cloud infrastructure (AWS/EKS) with infrastructure-as-code

Own incident response — in a trading system, every minute of downtime is real dollars at risk

Health monitoring for the agent fleet: which agents are scanning, which are stuck, which have the midnight rollover bug

Requirements
MANDATORY

  • Strong backend engineering — you write production code daily in at least two of: Go, Python, Node.js/TypeScript. Go being the preferred language for the runtime services
  • Experience building backend services from scratch at a startup: APIs, job scheduling, state management, distributed systems
  • Solid understanding of real-time systems where latency matters: websocket connections, condition-based evaluation, sub-second response requirements
  • Production experience with Postgres, Redis, and at least one analytics DB (ClickHouse, TimescaleDB, BigQuery)
  • Kubernetes experience — deploying, scaling, and debugging production workloads on AWS EKS
  • You’ve owned a system end-to-end: designed it, built it, deployed it, operated it, fixed it at 3am


STRONG PLUS

  • Experience with model serving / LLM infrastructure — deploying, scaling, and optimizing inference (vLLM, TGI, TensorRT-LLM, or managed endpoints)
  • Background in trading systems, exchange APIs, or fintech where uptime has direct financial consequences
  • Experience with onchain infrastructure: wallet operations, RPC nodes, transaction monitoring, DEX integration
  • Familiarity with MCP (Model Context Protocol) or similar agent-to-tool connectivity patterns
  • Experience building multi-agent platforms — orchestrating many independent processes sharing infrastructure but operating autonomously
  • Experience with CI/CD for systems where “deploy” means updating live trading agents, not just web servers


Benefits
  • Contract : Permanent role - Remote EU or US 
  • Benefits: healthcare and life insurance, retirement plan, sponsored transportation, gym, lunch card, home office equipment, and other perks.

Recruitment process :
  • Video call with CEO/CTO (30mn).
  • Take home coding test 
  • Code focus interview with tech teams
  • Final Interview​


Similar Jobs

Yesterday
Easy Apply
Remote
U.S.
Easy Apply
187K-278K Annually
Senior level
187K-278K Annually
Senior level
Artificial Intelligence • Enterprise Web • Software • Design • Generative AI
Build and operate reliable, scalable cloud infrastructure supporting Webflow products. Lead initiatives that improve reliability, reduce incident and triage load, and optimize cost and performance across multi-region environments. Develop infrastructure automation, modernize services, improve on-call and incident response, collaborate with product engineering teams, participate in design reviews, and mentor junior engineers. The role requires expertise in cloud platforms, containers, infrastructure as code, Cloudflare tooling, distributed systems, and software development.
Top Skills: AksAWSCloudflare WorkersDockerEcsEksGCPGkeGoKubernetesMesosNode.jsOpenshiftPulumiTerraformTypescript
Yesterday
Easy Apply
Remote
U.S.
Easy Apply
205K-316K Annually
Senior level
205K-316K Annually
Senior level
Artificial Intelligence • Enterprise Web • Software • Design • Generative AI
Lead the technical vision and architecture for backend systems supporting experimentation, personalization, analytics, and conversion optimization. Design and operate high-throughput distributed systems and APIs, build data infrastructure, and productionize AI/ML capabilities. Drive complex multi-team initiatives, influence engineering strategy, establish technical standards and operational excellence, and mentor senior engineers. Partner with product, data science, ML, and engineering teams to translate business needs into scalable solutions.
Top Skills: Ai/MlAPIsDistributed SystemsGoJavaPythonReal-Time Data Infrastructure
Yesterday
Easy Apply
Remote or Hybrid
Easy Apply
180K-270K Annually
Senior level
180K-270K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Staff Software Engineer leading insurance and cost-transparency features for a healthcare marketplace. Responsibilities include technical leadership, system design, roadmap ownership, hands-on development with C#, React, and AWS, mentoring engineers, improving engineering practices, collaborating cross-functionally, and delivering scalable, highly available products. The role requires strong product judgment, distributed backend systems expertise, developer productivity improvements, and experience working with third-party vendors and data providers.
Top Skills: AWSC#GenaiReact

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account