Your First Agent in Production
Part 8 / Production WorkflowsPre-Flight Checklist
Deploying your first agent to production is a milestone that most teams overthink or underthink. The overthinking camp spends months building infrastructure before running a single agent task. The underthinking camp gives an agent full access to their codebase on day one with no guardrails. Both approaches fail. The right approach is a controlled first deployment with clear boundaries, monitoring, and a plan for iteration.
Before deploying any agent to production, verify.
Workflow: Setting up your first production agent
Estimated time: 2-4 hours Prerequisites: Git repository, CI/CD pipeline, monitoring infrastructure
Step 1: Prepare the repository (30 minutes)
Create an AGENTS.md with your project overview, setup commands, code
conventions, testing requirements, and security notes (see Section 13.2
for the full template).
Then create the agent configuration:
# .agent/config.yaml
agent:
model: claude-sonnet-4.6
max_tokens_per_session: 200000
max_cost_per_session: 3.00
timeout_minutes: 30
permissions:
filesystem:
network:
commands:
observability:
tracing: true
cost_tracking: true
audit_logging: true
Step 2: Configure MCP servers (30 minutes)
Connect the agent to the tools it needs. At minimum, configure MCP servers for your filesystem (so the agent can read and write code), your version control system (so it can create branches and commits), and your test runner (so it can execute tests and read results). If your project uses a database, add a read-only database MCP server using the BFMCP pattern (Chapter 11) - expose high-level query operations, not raw SQL.
Start with 3-5 MCP servers. More than 10 creates tool definition bloat that wastes tokens on every turn (Chapter 5 covers Meta-MCP for compressing tool definitions). You can always add more servers later as the agent’s scope expands.
Verify each MCP server works independently before connecting it to the agent. Call each tool manually, confirm the response format, and check that error cases return structured error messages the agent can interpret.
Step 3: Set up observability (45 minutes)
You cannot debug agent failures without traces. Set up OpenTelemetry (Chapter 14) to capture every agent session as a trace, with spans for each LLM call, tool call, and decision point.
Configure your observability backend - Langfuse, Arize Phoenix, or Grafana Tempo with the OpenTelemetry exporter. At minimum, every trace should capture:
- Session metadata: session ID, task description, initiating user
- LLM calls: model name, input/output token counts, latency, cost
- Tool calls: tool name, arguments, response size, success/failure
- Cost: cumulative cost per session, with per-call breakdown
Set up two alerts from day one: one for sessions exceeding your cost cap ($3 default), and one for sessions exceeding your duration limit (30 minutes default). These catch runaway agents before they become expensive.
Step 4: Configure CI/CD integration (30 minutes)
Agent-generated code must pass the same quality gates as human-written code. Configure your CI pipeline to run the full backpressure stack (Chapter 32) on every agent-created pull request:
- Type checking -
tsc --noEmit,mypy, orgo vet - Linting - ESLint, Ruff, or golangci-lint with strict rules
- Tests - run affected tests, not the full suite (use
--changedSinceor pytest-testmon) - Security scanning - Semgrep or CodeQL for SAST
Configure the agent to read CI failure output so it can self-correct. The agent should be able to push a fix and re-trigger CI without human intervention. Set a maximum of 3 CI retry cycles - if the agent can’t pass CI after 3 attempts, escalate to a human with the full trace.
Step 5: First run (30 minutes)
Pick a low-risk task for the first run: a well-defined bug fix with a failing test, a documentation update, or adding unit tests for an untested function. Avoid database migrations, authentication changes, or public API modifications.
Write a clear task specification: what the agent should do, what files it should touch, what tests should pass when it’s done, and what it should not change. Vague specifications produce vague output.
Run the agent and watch the trace in real-time. Don’t intervene unless the agent is clearly stuck or approaching the cost cap. The goal of the first run is to observe the agent’s behavior, not to get a perfect result. Note where the agent hesitates, makes wrong assumptions, or lacks context - these observations will inform your AGENTS.md updates.
Step 6: Monitor and iterate (ongoing)
After the first successful run:
- Review the agent trace in your observability dashboard
- Check cost and token usage
- Identify any context gaps (update AGENTS.md)
- Gradually expand to more complex tasks
Workflow: Setting up Distill for monorepo context
Estimated time: 30 minutes
Monorepos present a unique context challenge. A typical monorepo has hundreds of packages, thousands of files, and millions of lines of code. No agent can process all of it. Distill solves this by building a compressed, deduplicated context index that agents can query for relevant information without loading the entire codebase.
The setup process is straightforward: install Distill, point it at your monorepo root, configure which directories to index (source code, documentation, ADRs) and which to skip (node_modules, build artifacts, generated code), and run the initial indexing. The index builds in minutes for most monorepos and updates incrementally as files change.
Once indexed, Distill exposes an MCP server that agents can query. Instead of loading entire files, agents ask Distill for the relevant context - “show me how error handling works in the payments service” - and receive a curated, deduplicated response that fits within their token budget.
Workflow: Implementing agent authorization with OpenFGA
Estimated time: 3-4 hours
Agent authorization with OpenFGA follows three steps. First, define your authorization model - the types (agent, repository, file, tool), the relationships (can_read, can_write, can_execute), and the rules (an agent that can_read a repository can_read all files in that repository). Second, populate the relationship tuples - which agents have which relationships to which resources. Third, integrate permission checks into your agent framework - before every tool call, check whether the agent has the required permission.
The most common mistake is making the authorization model too granular. You don’t need per-file permissions on day one. Start with per-repository permissions (this agent can access this repository) and per-tool-category permissions (this agent can use filesystem tools but not database tools). Add granularity as your needs become clear.
Related Concepts: All previous chapters Related Practices: Security Checklist