Posts tagged "operations"
-
Ops series recap: agent platforms are not that different from classic operations
AI platform operations recap: a table of twelve disciplines showing what carries over from DevOps and SRE and what is genuinely new for agents.
-
Ops series: capacity, rate limits and runaway agents
LLM rate limiting for agents: nested run, agent, tenant and provider limits stop one looping agent from starving the rest. With a config sketch and alerts.
-
Ops series: how to deploy AI agents with canaries and rollbacks
Deploy AI agents like any service: send a new model, prompt or skill to a small share of runs first, let evals and SLOs decide and keep a rollback ready.
-
Ops series: LLM FinOps for agents, attributing every token
LLM FinOps applies the cloud playbook to agents: tag every run, attribute spend per key and team, set budgets, alert early and export usage to cost reports.
-
Ops series: audit trails are not logs
An AI audit trail must be complete, attributable and tamper-evident; logs are for debugging. How to separate them on agent platforms and what goes where.
-
Self-hosted models for agents: what changes when inference is yours
Self-hosted LLM for agents: prompts stay in house, but you own capacity, updates and evaluation. Local runtimes versus batched servers and the ops work.
-
Ops series: turning runbooks into skills without losing control
Runbook automation AI done safely: turn a runbook into a skill plus narrowly scoped tools with approvals. A step-by-step conversion guide for ops teams.
-
Ops series: AI supply chain security and patching for agents
AI supply chain security for agents: inventory and patch models, MCP servers, skills, SDKs and images. A checklist for pinned, reviewed artefacts.
-
Ops series: AI incident response when the actor is an agent
AI incident response keeps the classic phases but needs a kill switch, tool revocation and policy freeze in seconds. A playbook for agent incidents.
-
Ops series: SLOs for agents, from task success to time to approval
Define an SLO for AI agents like for any service: task success rate, consistency over trials, latency and cost per task, plus an error budget for rollouts.
-
Ops series: observability for agents with traces, not transcripts
AI agent observability works best as distributed tracing: model and tool calls become spans with tokens, cost and decisions. Use OpenTelemetry, avoid lock-in.
-
Ops series: change management and prompt versioning for agents
Prompt versioning and change management for AI agents: apply semantic versioning, review and rollback to prompts, skills and policies like code.
-
Ops series: secrets for agents never belong in the context window
Secrets management for AI agents in practice: keep credentials on the tool side, give the model only handles, and never pass user tokens through to downstream APIs.
-
Ops series: identity for agents, or who is acting on whose behalf
AI agent identity: each agent needs its own identity plus a record of whom it acts for. Delegation vs impersonation, audience-bound tokens, no shared API keys.
-
Running agents is operations: the old disciplines still apply
AI agent operations is not a new field: identity, least privilege, change management, SLOs, cost and audit all still apply. The opening of an operations series.