Posts tagged "evaluation"
-
Measuring whether agents help: outcomes over impressions
AI productivity measurement needs outcomes, not impressions: accepted changes, review time, rework and cost per accepted change, with a metric tree.
-
Prompt injection testing: check your agents against hijacking
Prompt injection testing pairs a legitimate task with an injected one and checks both outcomes. Build suites for your own tools and keep containment in place.
-
LLM as a judge: a useful grader with known biases
LLM as a judge scales evaluation but has position, verbosity and self-enhancement bias. Which grader to use and how to calibrate a model grader.
-
Ops series: SLOs for agents, from task success to time to approval
Define an SLO for AI agents like for any service: task success rate, consistency over trials, latency and cost per task, plus an error budget for rollouts.
-
Quality gates for agent output: a checklist against slop
AI code quality gates let machines reject what humans should never read. A checklist of gates before review: tests, evals, size limits and policy checks.
-
Agent evals 101: tasks, graders and transcripts
AI agent evaluation checks the outcome in the environment, not just the final message. Learn tasks, trials, graders, transcripts and pass@k vs pass^k.