Skill Details
Runtime
(claude-5-opus)
Example user message
[userMessage]: Evaluate whether my customer-support LLM improved after changing the system prompt. I have 800 historical test cases, with categories for billing, refunds, account access, and general support. The previous version scored 87.4% overall and the new version scored 89.1%. I want to know whether the change is actually better before releasing it.
Example agent response
Evaluation Goal
Determine whether the new system prompt produces a reliable and practically meaningful improvement without introducing unacceptable regressions.
Evidence Check
The new version is 1.7 percentage points higher overall:
89.1% − 87.4% = +1.7 percentage points
However, the aggregate scores alone are insufficient to conclude that the new version is better.
Before making a release decision, verify:
Per-category results for billing, refunds, account access, and general support
Critical failure rates
Test-set composition and representativeness
Whether both versions were evaluated on the same 800 cases
Evaluation criteria and scoring method
Evaluator/judge configuration and calibration
Sample-size and uncertainty analysis
Any changes to model, retrieval, decoding, tools, or other system components
Regression Analysis
Do not invent category-level results from the overall scores.
Compare the two versions across each category and inspect whether any important subgroup or failure mode regressed.
If the same cases were evaluated by both versions, analyze the paired outcomes where the evaluation design supports it rather than treating the two aggregate percentages as independent measurements.
Reliability Check
Validate the evaluator against a representative human-reviewed subset when model-based judging is used. Inspect evaluator disagreements and determine whether the observed 1.7-point difference is statistically and practically meaningful.
Decision
MORE EVIDENCE REQUIRED
The new version has an observed overall improvement of 1.7 percentage points, but the supplied evidence does not establish that the improvement is reliable or that no unacceptable regression exists.
Next Action
Provide the category-level results, critical failure data, evaluation configuration, and raw or case-level comparison where available. Then perform the regression, uncertainty, and failure analysis before making the release decision.
Do not declare the new version better solely from the 89.1% versus 87.4% aggregate scores.
Get this skill + 9 more with Select$14 first month, then $19/mo · Cancel anytime$14 first month · then $19/mo
View SelectBy purchasing this skill, you agree to our terms of service
CLAUDE-5-OPUS
Design reliable LLM evaluation systems that define measurable criteria, build representative test sets, detect regressions, analyze failure patterns, and validate model behavior at scale across prompts, datasets, models, and versions. Apply structured metrics, human and model-based evaluation, statistical checks, reproducible test cases, traceability, and change-impact analysis to turn evaluation results into evidence-backed engineering decisions that remain useful as systems evolve.
...more
Added 4 days ago
