Loading...
Loading...
Found 2 Skills
Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates.
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.