Product.ai Logo

Product.ai

Machine Learning Engineer, Verification Engine

Posted 2 Days Ago
In-Office
Metropolitan, CA
260K-500K Annually
Senior level
In-Office
Metropolitan, CA
260K-500K Annually
Senior level
Design and own the verification engine that automates product truth verdicts: build confidence-scoring pipelines, adversarially robust models, a human-graded golden exam, and arbitration rulebooks. Raise machine-settled verdicts from 40% to 80% while maintaining 95%+ accuracy, embed eval-driven development, and scale solutions to millions of claims in production.
The summary above was generated by AI
You'd build the oracle the model can't game - the verification an entire truth layer ships through.

Product.ai is the verified truth layer for shopping - the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes is its first proof at scale: the code verification service - codes that actually work - at around $22M in revenue and roughly 60% margins. Founder-owned and profitable since 2009. No outside investors. No board. A small team - fewer than twenty operators - outbuilding companies 10x our size.

Strong people find us and keep finding us - they apply over months and years, because the field moves fast and the exact profile we need moves with it.

Why This Role Exists

Everything we sell rests on one question: is the claim true? The verification engine answers it at scale - machine verdicts that decide what we publish about a product, a price, a code. Today machines settle 40% of those verdicts and humans settle the rest. The constraint on growth is not more claims; it is whether the machine verdicts stay right as the volume climbs. This seat owns that - the scoring science behind every automated verdict, working directly with the founder, no layers between your models and the company's core asset.

The target is 80% of verdicts settled by machine while holding 95%+ accuracy against a human-graded exam set the model cannot see or game. You co-own the arbitration rulebooks - the rules that settle a contested verdict - with our truth scientist, so the layer an entire company ships through never rests on a single owner's judgment. It is also the most durable engineering seat in the building. Every model generation absorbs another layer of surface skill, so surface skills depreciate in months. Grading a verdict no human can check faster than the machine is an open frontier, and it compounds for years.

The System You'll Need to Model

  • Adversarial truth. Merchants, bots, and models all have incentives to game a verdict. A merchant wants its expired code to read as live; a model grades its own family roughly 30 points too kind; a bot wants to look like a shopper. Your verifier has to be structurally harder to fool than any of them, because every party at the table is motivated to fool it.
  • The confidence-scoring pipeline. A verdict is not a yes or a no; it carries a confidence and a decay. A price is stale in days, a materials fact holds for years. The pipeline has to score how sure the machine is and stay honest when it is not sure - a confidently wrong verdict is worse than an abstention.
  • The golden-set exam - grading the grader. The hardest problem here: how do you measure a grader when the thing it grades is a judgment no human can make faster than the machine? The answer is a human-graded exam set the model cannot see, cannot train on, and cannot game. Designing that exam is the science.
  • Eval-driven development as the company's core loop. Evals are not a testing afterthought here; they are how engineering works. Agent loops write the code and the content; your corpora and gates decide whether a loop's output ships or gets rejected. The eval is the spec.
  • Cortex - the shared AI brain. You work inside Cortex, the governed AI substrate that runs the company and is the product family we sell. Every operator works through governed AI sessions, and the substrate answers its own questions from 8,600+ documents. Nobody else runs the company on the machine they sell - and your verification is what keeps its answers trustworthy.
  • Scale. Hundreds of thousands of merchants, millions of claims. A verification method that works on a hundred cases and falls over at a million is not a method; it is a demo. Everything you build has to hold at the volume the business actually runs at.


If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.

What You Will Own

  • The scoring science. The machine verdicts that decide what we claim is true about shopping - the models, the thresholds, and the rulebooks behind every automated verdict. The number you own: 40% of verdicts settled by machine today, taken to 80%, at 95%+ accuracy against the human-graded exam. Like every outcome here, it is falsifiable - it ships with a test a stranger could run.
  • The confidence-scoring pipeline. How sure the machine is, per claim class, with the decay each class carries. You own its calibration and its honesty: a verdict that claims 95% and is right 70% of the time is a bug, and it is yours.
  • The golden-set exam. The human-graded exam the model cannot see or game. You grow it, you keep it honest, and you give it teeth. When a verdict is disputed, the exam decides - not the loudest engineer in the room.
  • The arbitration rulebooks, co-owned. The rules that settle a contested verdict, authored with our truth scientist. Two owners by design: the truth an entire company ships through does not rest on one person.
  • Your seat charter. Within your first quarter you co-sign a charter for this seat - one machine-checkable number that proves it is working, and a written split of what you decide freely versus what you bring to the founder to decide. This is the model we run: a real authority split, in writing, not a job description.


Visibility here is registered architecture decisions and outcome movement - not hours, meetings, or activity.

Who You Are

You interrogate every green checkmark. A passing eval is a claim, and claims get challenged - you ask what the test could not have caught before you ask what it confirmed. You form working models of complex systems on your own, notice where your model is wrong, and update fast. You write clearly, because clear writing is evidence of clear thought.

You treat agents as leverage you verify, never as an oracle you trust. You can do this job by hand and prove it - hand-grade a judge, hand-build a corpus, hand-verify a claim set - and that mastery is exactly what lets you trust, or reject, the verdict an agent hands back. You move between a confidence-tier scheme in the morning and the verifier fleet that enforces it by evening without getting stuck at either altitude. The expensive thing here is a redo cycle, never the compute.

You have probably built an eval harness another team came to depend on, a regression corpus that caught a real failure before users did, an LLM-judge pipeline where you measured the judge's bias instead of trusting it, or a test gate for a system whose output is never the same twice. We care about the artifact and the reasoning behind it far more than where you did it.

Who this isn't for. This role is wrong if you optimize for leaderboard scores or trust vendor-reported evals - the work here is deciding what a score even means, not chasing one. It is wrong if your instinct when output looks weak is to massage the wording rather than redesign the verification, and wrong if you would rather publish a finding than ship a gate. It is wrong if you want tightly-scoped tickets and a lane to stay in; the scope of this seat is the whole verification surface. And it is wrong if you are comfortable shipping what an agent produced without being able to say why it is right, or letting an agent grade its own work. You'll be happiest here if you have high agency, think in corpora, and want your verification to be the reason an entire truth layer can be trusted.

How We Evaluate

We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.

  • Async video screen. Brief and on your own time: 5-6 questions, about 15 minutes. We want to see how you think, not how you present.
  • Calls with company stakeholders. Short conversations with key members of the team.
  • Conversation with the founder. How you model the system above, where you push back, and whether you can hold the argument live.
  • Paid work trial. One week of paid, real verification work in our real environment - live loops, live corpora, real verdicts. It is paid because your time is worth paying for, and we say so up front because high-agency people prefer real stakes to a take-home. We watch how you get grounded, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest.


  • If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the artifact and the reasoning, not the pedigree.

    Compensation & Ownership

    Total first-year comp: $400,000 - $500,000 (base + ownership + profit sharing). Base: $260,000 - $330,000 - top of market for machine learning engineering.

    Ownership is real and liquid. Profits Interest Units (PIUs) - Class B Membership Interests at a $0 strike, real ownership from day one, with capital-gains treatment. Annual pro-rata profit sharing from free cash flow - real cash every year, not a promise tied to a distant exit. An annual tender offer buys back vested interests at fair market value, so you can turn ownership into cash every year without waiting for an IPO. 100% premium coverage for you and your family. The token budget is effectively unlimited, steered by ROI and never capped - high usage is encouraged, because the harness tracks what the spend returned.

    This structure is built to mint partners. When the company wins, you win - in real, liquid dollars, every year.

    Based in Santa Monica, Los Angeles - in person, five days a week. The rooms are real rooms. Relocation support available for the right builder.
    HQ

    Product.ai Los Angeles, California, USA Office

    Our office is centrally located at the intersection of Santa Monica and Brentwood on a trendy section of Wilshire. Offering expansive views of the ocean to downtown LA, our high rise building sits right next to some of LA's most popular restaurants, cafes, juice bars and brunch spots.

    Similar Jobs at Product.ai

    2 Days Ago
    In-Office
    250K-475K Annually
    Senior level
    250K-475K Annually
    Senior level
    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    Own and scale an end-to-end fleet of checkout robots: browser automation, merchant-family classification, and scheduling. Improve machine-tested checkout coverage from ~21% to >80% while controlling cost per check. Navigate anti-bot defenses, instrument correctness with dashboards and ledgers, and co-author the seat charter that defines measurable coverage and reliability. Ship fast, diagnose hard failures, and operate production fleets with strong judgment and ownership.
    Top Skills: Browser AutomationCrawling SystemsHeadless ChromeNode.jsPlaywrightProxy RotationPythonQueue-Backed Job SystemsWeb Scraping
    2 Days Ago
    In-Office
    199K-500K Annually
    Expert/Leader
    199K-500K Annually
    Expert/Leader
    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    Own and monetize the agent-economy revenue line: price and sell keyed API access, convert free merchant alerts into paid find-and-fix contracts, manage affiliate-network relationships, and build platform partnerships and pilots (ChatGPT apps, App Intents, Gemini, Claude). Deliver a measurable first-quarter charter and signed pilots; run end-to-end commercial motion from packaging and pricing to partner integration.
    Top Skills: APIsApple App IntentsChatgptClaudeCortexGeminiKeyed ApiMcp
    2 Days Ago
    In-Office
    199K-480K Annually
    Expert/Leader
    199K-480K Annually
    Expert/Leader
    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    Own and build the end-to-end billing spine and commission engine that converts a free API into paid revenue. Architect keyed developer access, metering, quotas, spend caps, attribution, affiliate ingestion/reconciliation, and agent-commerce billing. Ship two public cutovers, instrument pipelines to tie failures to dollars, and collaborate with commercial operators to deliver measurable revenue outcomes.

    What you need to know about the Los Angeles Tech Scene

    Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

    Key Facts About Los Angeles Tech

    • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
    • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
    • Key Industries: Artificial intelligence, adtech, media, software, game development
    • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
    • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
    • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account