Founding Engineer, RL & Evals

Runtime Labs builds Óra at ora.app.

As a Founding Engineer focused on RL & Evals, you will own how Runtime Labs measures model behavior in Óra: planning, scheduling, recommendations, memory use, tool use, and long-horizon interaction. You will define what good looks like, build the datasets and harnesses that score it, and close the loop so evaluation signal drives product decisions, model selection, and—when the data justifies it—post-training and reinforcement learning.

This is not generic chatbot evaluation. The domain is agentic behavior over time: whether a model uses the right context, makes good plans, chooses appropriate actions, recovers from mistakes, and improves as interaction data compounds.

About the Role

  • Own the eval roadmap for planning, scheduling, recommendations, memory, tool use, and long-horizon model behavior in Óra
  • Define task- and system-level metrics that separate useful behavior from superficially plausible outputs—including temporal reasoning, retrieval quality, action selection, constraint satisfaction, and recovery from failure
  • Build and version evaluation datasets from synthetic scenarios, curated examples, production failures, and longitudinal interaction; turn real failures into durable regression cases
  • Design dataset schemas that preserve context, expected behavior, provenance, annotations, and outcomes—with clear conventions for sampling, contamination control, and reproducible comparisons
  • Build developer-friendly harnesses for writing, running, inspecting, and comparing evals; integrate them into model experimentation and everyday product engineering
  • Support offline replay and controlled comparison across models, prompts, retrieval strategies, and agent policies; make failures traceable from interaction through retrieval, reasoning, tool calls, and final action
  • Establish release criteria so model and product changes are measured before they reach users
  • Close the loop from eval signal to shipping decision: prioritize regressions, guide prompting/retrieval/architecture choices, and develop preference or reward signals from evaluation, user behavior, and longitudinal outcomes when warranted
  • Explore post-training and reinforcement-learning approaches when evaluation quality and interaction data justify them
  • Partner with Product, Infrastructure, and Data Infrastructure so production interaction becomes structured learning signal

About the Person

  • Research-minded engineer who cares deeply about measurement, experimental rigor, and why model behavior changes
  • Strong ML foundation with experience evaluating, training, or improving language-model or agentic systems
  • Moves between research questions and production engineering: define the experiment, build the dataset, implement the harness, analyze failures, ship the change
  • Strong intuition for evaluation design and the failure modes of automated metrics, model-based graders, and human labels
  • Reasons precisely about noisy signals, baselines, distributions, regressions, and experimental validity
  • Strong software engineer who builds durable evaluation infrastructure—not one-off notebooks or manual review alone
  • Interested in reinforcement learning, preference optimization, post-training, agent evaluation, and learning from interaction
  • Excited by planning, temporal reasoning, recommendations, memory, and models that act over long-lived user context
  • Wants high ownership as an early long-term partner defining how Runtime Labs measures and improves intelligent behavior

Experience

  • Experience building evaluation systems for LLMs, agents, recommendations, search, ranking, or decision-making systems
  • Experience with reinforcement learning, preference optimization, reward modeling, or post-training
  • Experience building datasets, annotation systems, graders, replay infrastructure, or experimentation platforms
  • Familiarity with model tracing, tool-use evaluation, retrieval evaluation, and multi-step agent benchmarks
  • Experience turning production model failures into reproducible test cases and measurable improvements
  • Background in sequential decision-making, recommendation systems, temporal reasoning, or human-AI interaction

Location & Compensation

Location: San Francisco. We are building the founding team in person and expect to work closely together during the company's early formation.

Salary: $150k–$200k depending on experience and role.

Equity: meaningful early-stage ownership for long-term partners.

Exceptional candidates may be considered outside this range.

We encourage you to apply even if you do not meet every qualification.

Apply for this position

← All positions