Posts tagged "quality"
-
Prompt hygiene: a checklist for teams drowning in prompts
Prompt management for teams: a checklist covering owner, version, eval, scope and review, plus the rule to move repeated know-how into skills.
-
LLM as a judge: a useful grader with known biases
LLM as a judge scales evaluation but has position, verbosity and self-enhancement bias. Which grader to use and how to calibrate a model grader.
-
Quality gates for agent output: a checklist against slop
AI code quality gates let machines reject what humans should never read. A checklist of gates before review: tests, evals, size limits and policy checks.
-
Small verified steps beat one big prompt: incremental agent workflows
Incremental agent workflows break work into small steps with acceptance checks, progress notes and a clean state after each one. Here is the pattern.
-
Agent evals 101: tasks, graders and transcripts
AI agent evaluation checks the outcome in the environment, not just the final message. Learn tasks, trials, graders, transcripts and pass@k vs pass^k.
-
Too many prompts: how prompt sprawl turns into AI slop
AI slop is output that grows faster than anyone can review it. How prompt sprawl lowers quality, and why fewer, smaller, verified steps beat more prompts.