Posts tagged "explainer"
-
Beyond tools: MCP sampling, elicitation and tasks
MCP sampling lets servers request model calls; elicitation asks users for input; tasks track long work. Each moves control and needs its own policy.
-
An AI governance operating model for agents: who owns what
An AI governance operating model fails when nobody owns the agent registry, policies or budget. Roles and a RACI for platform, security, owners and reviewers.
-
Long-running AI agents: progress files, clean states, checkpoints
Long-running AI agents lose context between sessions. Keep state in a feature list, progress notes and git commits to make runs resumable and auditable.
-
LLM as a judge: a useful grader with known biases
LLM as a judge scales evaluation but has position, verbosity and self-enhancement bias. Which grader to use and how to calibrate a model grader.
-
Agent evals 101: tasks, graders and transcripts
AI agent evaluation checks the outcome in the environment, not just the final message. Learn tasks, trials, graders, transcripts and pass@k vs pass^k.
-
Context engineering: the smallest set of tokens that does the job
Context engineering treats the context window as a finite budget. Practical ways to keep it small: compaction, sub-tasks with own context, notes.
-
What is the Model Context Protocol? A plain explanation for platform teams
What is MCP? The Model Context Protocol standardises how an agent host finds and calls tools, resources and prompts on servers. Plain guide for platform teams