C

Clera

remote

Data Scientist — Agent Evaluations & Quality

Full Time今天
Engineering

职位描述

About the Role

This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a Data Scientist — Agent Evaluations & Quality, you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality.

This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality — ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.

  • Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.

  • Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement.

  • Compare models, prompts, and implementations using rigorous offline experiments and production evidence.

  • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.

  • Turn production failures into regression cases and continuously close gaps in evaluation coverage.

  • Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.

  • Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions.

What We're Looking For

Required

  • 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.

  • Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.

  • Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets.

  • Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems.

  • Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.

  • Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes.

  • Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers.

  • Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.

Nice to Have

  • Hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms.

  • Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software.

  • Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems.

What makes you a great fit

  • You're product-oriented — you prioritize metrics tied to real user outcomes, not just convenient measurements.

  • You drive ambiguous quality questions from evaluation design all the way into product decisions.

  • You write maintainable, production-quality code — not just ad-hoc notebooks.

  • You collaborate naturally with engineers and are comfortable digging into traces and system internals.

Location

This role is on-site. Visa sponsorship is not available for this position.

Compensation & Benefits

Compensation details were not provided for this listing. A competitive package commensurate with experience is expected at this stage of company growth.

Find more English Speaking Jobs in United Kingdom on Arbeitnow

Applying for Data Scientist — Agent Evaluations & Quality?

Build a resume tailored to this role at Clera in minutes — free, and formatted to pass applicant tracking systems.

如何申请

准备好申请了吗?你需要准备:

  • 一份突出相关经验的最新简历
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

关于 Clera

C

Clera

remote

常见问题

Where is the Data Scientist — Agent Evaluations & Quality position at Clera located?

The Data Scientist — Agent Evaluations & Quality role at Clera is based in remote.

What type of employment is the Data Scientist — Agent Evaluations & Quality role at Clera?

This Data Scientist — Agent Evaluations & Quality position is offered as Full Time.

What skills are needed for the Data Scientist — Agent Evaluations & Quality role at Clera?

Key skills and focus areas for this role include Engineering.

How do I apply for the Data Scientist — Agent Evaluations & Quality position at Clera?

You can apply for the Data Scientist — Agent Evaluations & Quality role at Clera directly from this page. Create a professional, ATS-ready resume with Clever CV to strengthen your application before you apply.

探索更多机会

来自知名企业的数千个职位正在等你