Posts tagged "how-to"
-
Building an internal skills catalog: review, ownership and discovery
Build a skills catalog for your organisation: submit, review, publish with owner and version, then deprecate. Borrows ideas from public registries.
-
Ops series: how to deploy AI agents with canaries and rollbacks
Deploy AI agents like any service: send a new model, prompt or skill to a small share of runs first, let evals and SLOs decide and keep a rollback ready.
-
Structured output for agents: schemas, contracts and ownership
Structured output agents hand over typed data, not free text. Define JSON schemas for tool results and outputs, version them and name an owner per contract.
-
Ops series: LLM FinOps for agents, attributing every token
LLM FinOps applies the cloud playbook to agents: tag every run, attribute spend per key and team, set budgets, alert early and export usage to cost reports.
-
Enterprise AI agent rollout: pilot, baseline, then scale
An enterprise AI agent rollout works best with a small pilot, telemetry from day one, spend limits and written exit criteria for each wider step.
-
Self-hosted models for agents: what changes when inference is yours
Self-hosted LLM for agents: prompts stay in house, but you own capacity, updates and evaluation. Local runtimes versus batched servers and the ops work.
-
Ops series: turning runbooks into skills without losing control
Runbook automation AI done safely: turn a runbook into a skill plus narrowly scoped tools with approvals. A step-by-step conversion guide for ops teams.
-
LLM cost optimization for agents: caching, code execution and budgets
LLM cost optimization for agents: where prompt caching pays off, when code execution beats tool calls, and why hard per-run token budgets protect security too.
-
Claude Code hooks as policy enforcement points for coding agents
Claude Code hooks run code before a tool call, so a PreToolUse hook can enforce policy locally. How to write one, where it stops and why a central gate stays.
-
AI agent sandboxing: filesystem and network isolation together
AI agent sandboxing works only when file access and network egress are both limited. Why one boundary is not enough and how allow-lists cut approvals.
-
Versioning agent skills like any other dependency
Skill versioning treats agent skills as dependencies: semantic versions, changelog, pinning and evals that guard upgrades. What counts as a breaking change.
-
Draw the data flow before you write the first prompt
An agent data flow diagram shows where untrusted data enters, where private data lives and where anything leaves. One-page template for trifecta risks.
-
One MCP server per system: MCP server design you can govern
MCP server design rules for governable agents. One small server per system, namespaced tools, and read/write separation instead of one giant gateway.
-
Small verified steps beat one big prompt: incremental agent workflows
Incremental agent workflows break work into small steps with acceptance checks, progress notes and a clean state after each one. Here is the pattern.
-
Anatomy of an agent skill: SKILL.md, progressive disclosure and scope
SKILL.md is the entry point of an agent skill: its short description decides when the rest loads. Progressive disclosure, a minimal skill and scoping rules.
-
Writing tools agents can actually use: names, schemas and errors
Tool design for AI agents is interface design for a model reader. A checklist for names, namespaces, narrow schemas and errors, and how small tools ease policy.