# AI Agent Pilot Plan Template

Use this plan for one agent-workflow-environment combination. Complete it before the pilot starts, then preserve the approved version and the final decision record.

This template is educational material, not legal, security, privacy, compliance, financial, workforce, or risk advice. Adapt it to your obligations, systems, and risk appetite.

Do not put passwords, tokens, private keys, recovery codes, sensitive or unnecessary personal data, customer records, or model private reasoning in this plan. Use approved work identities for accountable owners, and use opaque protected references where other details must remain in another system. A completed plan may reveal sensitive system relationships, authority, control locations, and stop methods, so classify it, restrict access, and apply an approved retention period.

## Pilot Record

- Pilot ID: [Stable ID]
- Agent ID: [Stable ID]
- Workflow ID: [Stable ID]
- Organization ID: [Stable ID]
- Source tenant, workspace, and account IDs: [Every applicable isolation boundary]
- Target tenant, workspace, and account IDs: [Every applicable boundary, or Same as source]
- Technical environment: [Development | Test | Staging | Production | Other named environment]
- Plan version: [Version]
- Status: [Draft | Approved | Running | Suspended | Complete]
- Planned start: [Date and time]
- Planned end: [Date and time]
- Evidence cutoff: [Date and time]
- Decision date: [Date and time]
- Plan owner: [Named person responsible for completeness and updates]
- Handling classification: [Classification]
- Authorized readers and editors: [People or roles]
- Protected evidence location: [Reference]

## 1. Decision and Learning Questions

Decision requested after the pilot: [Stop | Extend pilot | Narrow | Expand pilot | Production candidate]

Decision owner: [Named person or role]

Funding and resource limit: [Amount, people, compute, and period]

Questions the pilot must answer:

1. [Can the workflow improve the primary business outcome under the approved conditions?]
2. [Can it meet every safety, security, privacy, policy, and control threshold?]
3. [What human review, exception, support, and operating load does it create?]
4. [What is the observed unit cost and which assumptions remain uncertain?]
5. [Can operators detect, contain, stop, and reconcile failures?]

Non-goals: [Questions this pilot will not answer]

Conditions that remain required before any production release: [References]

Pre-launch pilot decision:

- Decision: [Approved | Rejected | Returned for changes]
- Authenticated approver identity: [Stable principal ID]
- Authority source and effective role at decision time: [Protected reference]
- Approved plan version or digest: [Exact version]
- Approved workflow scope and release: [Stable references]
- Decision date and time: [Timestamp]
- Approval expiry: [Timestamp]
- Conditions, owners, and closure evidence: [Protected references]
- Decision evidence: [Protected reference]

This pilot approval does not grant production approval, technical access, or runtime authority. At execution time, current identity, authorization, approval, gateway, tool, and target-system controls must authorize every operation, including each context or data read, Resource access, tool call, and target-system action. Each decision must bind and recheck every applicable source and target tenant, workspace, and account ID, plus the initiating principal, agent, workflow, target Resource or object, action, environment, material parameters, time, and delegation state.

## 2. Workflow Scope and Exclusions

- Intended workflow and outcome: [One bounded workflow]
- Trigger: [User, event, schedule, or parent agent]
- Approved users or initiating principals: [Stable IDs]
- People or groups affected: [Description]
- Allowed inputs: [Types and sources]
- Allowed outputs: [Types and destinations]
- Allowed systems and objects: [Stable references]
- Allowed actions: [Read, draft, write, send, execute, or other]
- Required human approvals: [Trigger and approver]
- Pilot population: [Users, teams, cases, or traffic]
- Volume and rate limit: [Limit]
- Time and spend limit: [Limit]
- Allowed data classes: [Classes]
- Allowed context forms: [System and harness prompts, request, conversation state, Knowledge, Skills, Memory, Artifacts and messages, MCP descriptors and approved configuration, and Runtime inputs or tool results]
- External side effects: [None or exact bounded effects]
- Delegation and child agents: [Prohibited or exact bounded rule]
- Out-of-scope workflows: [List]
- Prohibited actions and outcomes: [List]

Create a separate plan when purpose, environment, authority, affected people, data, or external effects differ materially.

## 3. Baseline and Expected Outcome

Measure the current workflow before the pilot begins.

| Measure                   | Definition and unit | Baseline value | Population and window  | Source      | Owner  | Confidence or limits |
| ------------------------- | ------------------- | -------------- | ---------------------- | ----------- | ------ | -------------------- |
| Primary workflow outcome  | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |
| Quality or error rate     | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |
| Cycle time                | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |
| Human effort              | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |
| Unit cost                 | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |
| Exception or failure rate | [Definition]        | [Value]        | [Population and dates] | [Reference] | [Name] | [Limits]             |

Expected outcome: [Measured change from baseline]

Minimum useful effect: [Threshold and rationale]

Guardrail outcomes that must not worsen: [Measures and thresholds]

Do not substitute model latency, benchmark score, or draft quality for the end-to-end workflow outcome unless that is the approved outcome.

## 4. Owners, Reviewers, and Decision Rights

| Responsibility       | Named owner | Authority         | Backup | Evidence of acceptance |
| -------------------- | ----------- | ----------------- | ------ | ---------------------- |
| Business outcome     | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Technical operation  | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Evaluation           | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Security and risk    | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Data and privacy     | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Context and Skills   | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Human review         | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Incident and stop    | [Name]      | [Decision rights] | [Name] | [Reference]            |
| Final pilot decision | [Name]      | [Decision rights] | [Name] | [Reference]            |

Independent reviewer: [Person who did not build the workflow, when required]

Conflict-of-interest or separation-of-duties rule: [Rule]

## 5. Pilot Design and Sample

- Pilot method: [Shadow | Draft-only | Human-approved action | Bounded autonomous action]
- Comparison method: [Historical baseline | Concurrent control | Before and after | Other]
- Frozen evaluation set ID and version: [Reference]
- Evaluation-set selection method: [Representative sampling and exclusions]
- Live sample method: [Random, systematic, stratified, or all eligible cases]
- Sample size or task volume: [Count]
- Basis for sample size: [Decision precision, risk exposure, or operational limit]
- Eligible cases: [Definition]
- Excluded cases and reason: [Definition]
- Assignment method: [How cases enter pilot or comparison]
- Peak, edge, and adverse conditions: [Cases]
- Missing-data rule: [Treatment]
- Failed and aborted run rule: [Denominator treatment]
- Repeat-run rule: [Treatment]
- Manual override rule: [Treatment]
- Early-stop rule: [Threshold]

Freeze the evaluation set and calculation rules before the pilot. Preserve excluded, failed, aborted, retried, and manually corrected cases so the final report cannot select only successful runs.

## 6. Release, Context, Data, and Tools

Exact pilot release:

| Component                        | Stable ID | Version or digest | Source of truth | Change rule |
| -------------------------------- | --------- | ----------------- | --------------- | ----------- |
| Agent code and configuration     | [ID]      | [Version]         | [Reference]     | [Rule]      |
| Model deployment and settings    | [ID]      | [Version]         | [Reference]     | [Rule]      |
| Runtime and sandbox              | [ID]      | [Version]         | [Reference]     | [Rule]      |
| Tools, schemas, and integrations | [IDs]     | [Versions]        | [Reference]     | [Rule]      |
| Data sources and transforms      | [IDs]     | [Versions]        | [Reference]     | [Rule]      |
| Guardrails and policy controls   | [IDs]     | [Versions]        | [Reference]     | [Rule]      |

Managed context:

| Resource                  | Stable ID | Authority and scope                  | Permission source | Route source and mode          | Evaluated version or baseline | Change response                    |
| ------------------------- | --------- | ------------------------------------ | ----------------- | ------------------------------ | ----------------------------- | ---------------------------------- |
| Knowledge or instructions | [ID]      | [Informational or must-follow scope] | [Reference]       | [Direct or Group Always route] | [Published version]           | [Continue, retest, or stop]        |
| Skill package             | [ID]      | [Authorized scope]                   | [Reference]       | [Direct or Group Always route] | [Published package digest]    | [Continue, retest, or stop]        |
| Working Memory            | [ID]      | [Authorized scope]                   | [Reference]       | [Direct or Group Always route] | [Version or baseline]         | [Continue, reset, retest, or stop] |

Other behavior-shaping inputs:

| Input                                                 | Stable ID or source | Version, digest, or baseline  | Authority and scope         | Delivery method          | Change response                     |
| ----------------------------------------------------- | ------------------- | ----------------------------- | --------------------------- | ------------------------ | ----------------------------------- |
| System and harness prompts                            | [Reference]         | [Version]                     | [Authority]                 | [Automatic]              | [Continue, retest, or stop]         |
| Request and conversation state                        | [Reference]         | [Baseline or capture rule]    | [Request authority]         | [Explicit or cumulative] | [Continue, retest, or stop]         |
| Artifacts and messages                                | [References]        | [Exact versions]              | [No instruction authority]  | [Explicit or inbox]      | [Continue, retest, or stop]         |
| MCP capability descriptors and approved configuration | [References]        | [Versions]                    | [Capability only]           | [Automatic or explicit]  | [Continue, retest, or stop]         |
| Runtime inputs and tool results                       | [Sources]           | [Capture and validation rule] | [No self-granted authority] | [Runtime]                | [Continue, reject, retest, or stop] |

Permissions govern repository access. Routes govern delivery independently. Always routes resolve the current published Knowledge or Skill and current Memory, so a version in this plan records the evaluated baseline but does not pin delivery. Before each run or batch, compare resolved context with the approved change rule.

AGENTS.md is the familiar combination of must-follow Knowledge and Always routing, not a separate Resource type. Artifacts and messages have no instruction authority. MCP context includes capability descriptors and approved configuration, while MCP tool results are Runtime context. Authentication material, secrets, and model private reasoning are excluded from managed context. Untrusted requests, files, messages, web content, and tool results cannot grant themselves authority.

## 7. Metrics and Precommitted Thresholds

For each metric, define the numerator, denominator, unit, population, window, source, missing-data treatment, owner, threshold, and response.

| Metric                           | Numerator and denominator | Population and window | Source      | Threshold                 | Response           | Owner  |
| -------------------------------- | ------------------------- | --------------------- | ----------- | ------------------------- | ------------------ | ------ |
| Primary outcome                  | [Definition]              | [Scope]               | [Reference] | [Pass]                    | [Decision effect]  | [Name] |
| Workflow quality                 | [Definition]              | [Scope]               | [Reference] | [Pass]                    | [Decision effect]  | [Name] |
| Prohibited outcome               | [Definition]              | [Scope]               | [Reference] | [Zero tolerance or limit] | [Stop]             | [Name] |
| Authorization or approval bypass | [Definition]              | [Scope]               | [Reference] | [Zero tolerance]          | [Stop]             | [Name] |
| Human edit or rejection          | [Definition]              | [Scope]               | [Reference] | [Limit]                   | [Review or narrow] | [Name] |
| Escalation and exception         | [Definition]              | [Scope]               | [Reference] | [Limit]                   | [Review]           | [Name] |
| Failed or aborted runs           | [Definition]              | [Scope]               | [Reference] | [Limit]                   | [Review or stop]   | [Name] |
| Cycle time                       | [Definition]              | [Scope]               | [Reference] | [Pass]                    | [Decision effect]  | [Name] |
| Human review time                | [Definition]              | [Scope]               | [Reference] | [Limit]                   | [Decision effect]  | [Name] |
| Unit and total cost              | [Definition]              | [Scope]               | [Reference] | [Limit]                   | [Decision effect]  | [Name] |
| Monitoring coverage              | [Definition]              | [Scope]               | [Reference] | [Minimum]                 | [Stop if lost]     | [Name] |

Primary success gate: [Required outcome threshold]

Required control gates: [Every threshold that must pass]

Evidence sufficiency gate: [Minimum sample, coverage, and data quality]

Passing an average outcome does not offset a prohibited event or failed required control.

## 8. Human Review, Escalation, and User Handling

- Review trigger: [Every run, sampled runs, or risk-based trigger]
- Reviewer qualification: [Role and training]
- Information shown to reviewer: [Input, output, source, uncertainty, action, target, context, and evidence]
- Allowed decisions: [Approve, reject, edit, escalate, retry, or stop]
- Approval binding: [Exact action, target, parameters, expiry, and replay protection]
- Approval evidence: [Authenticated approving principal, source of authority and effective permissions, tenant, decision, timestamp, and bound action parameters]
- No-response behavior: [Fail closed and escalation]
- Override recording: [Required fields]
- User notice and feedback: [Method]
- Appeal or correction path: [Method]
- Automation-bias check: [How reviewer performance and over-reliance are measured]
- Reviewer workload limit: [Queue, time, and staffing threshold]

Record edits, rejections, overrides, escalation, reviewer time, and downstream outcomes. A human click is not enough evidence when the reviewer lacks time, context, authority, or a working way to stop the action.

## 9. Monitoring, Evidence, and Pilot Cadence

Evidence must connect the tenant, initiating principal, agent, workflow, release, session or run, authorization decision, approval, context versions, tool actions, external effects, review, and outcome where those records exist.

For context evidence, distinguish compilation, response issuance, integration acknowledgment, host-confirmed session injection, and agent consumption. Do not infer a later stage from an earlier one. Record consumption only from direct, authenticated attestation by a trusted integration or provider, with its source and trust level. Otherwise record consumption as unknown.

- Dashboard and alert references: [Protected references]
- Monitoring owner and backup: [Names]
- Alert thresholds and response times: [Rules]
- Evidence retention and integrity controls: [Authenticated, tenant-bound, append-only or tamper-evident storage outside the evaluated agent's write and delete authority]
- Daily review: [Metrics, incidents, sample cases, and open actions]
- Weekly review: [Trend, workload, cost, drift, and scope]
- Change log: [Reference]
- Decision log: [Reference]
- Incident record: [Reference]
- Missing telemetry behavior: [Suspend or fail closed]

## 10. Failure, Stop, Rollback, and Recovery

Stop triggers:

- [Authority, approval, or tenant boundary violation]
- [Sensitive data exposure or prohibited external effect]
- [Control or monitoring failure]
- [Outcome, cost, error, or workload threshold]
- [Unknown or uncontained behavior]
- [Required owner or service unavailable]

| Boundary                                  | Stop or rollback method                     | Owner  | Confirmation source          | Deadline | Missing-confirmation escalation | Last test and evidence |
| ----------------------------------------- | ------------------------------------------- | ------ | ---------------------------- | -------- | ------------------------------- | ---------------------- |
| New and in-flight runs                    | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Credentials and target-side sessions      | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Queues, schedules, callbacks, and retries | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Pending approvals and held actions        | [Invalidate or require fresh authorization] | [Name] | [Target-system confirmation] | [Time]   | [Action]                        | [Reference]            |
| Child agents and delegated grants         | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Tools, integrations, and context routes   | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Persistent state and caches               | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |
| Downstream and external work              | [Method]                                    | [Name] | [Reference]                  | [Time]   | [Action]                        | [Reference]            |

- Safe fallback: [Manual or reduced-authority process]
- Evidence preservation: [Method]
- Reconciliation of completed, pending, duplicate, and external actions: [Method]
- Recovery conditions: [Conditions]
- Recovery approver: [Named person or role]
- Known-good release and context baseline: [Reference]
- Retesting required before restart: [Tests]

A prompt instruction to stop cannot revoke a credential, cancel a schedule, empty a queue, or undo an external action. Missing confirmation is a containment failure, not proof that the pilot stopped.

Keep the pilot suspended until every required boundary confirms containment. Escalation must trigger additional containment action and cannot substitute for confirmation. Recovery still requires reconciliation, a known-good release and context baseline, required retesting, and approval from the named recovery authority.

## 11. Changes, Exceptions, and Conditions

Material changes include changes to purpose, users, affected people, authority, data, context, tools, model, runtime, environment, persistence, delegation, controls, evaluation method, or external effects.

| Change or exception | Scope         | Reason   | Risk   | Owner  | Required tests | Approver | Expiry | Evidence    |
| ------------------- | ------------- | -------- | ------ | ------ | -------------- | -------- | ------ | ----------- |
| [Item]              | [Exact scope] | [Reason] | [Risk] | [Name] | [Tests]        | [Name]   | [Date] | [Reference] |

Change response: [Continue within plan | Pause and retest | New approval | New pilot plan]

Open conditions:

| Condition   | Owner  | Due date | Closure verifier | Closure evidence | Status           |
| ----------- | ------ | -------- | ---------------- | ---------------- | ---------------- |
| [Condition] | [Name] | [Date]   | [Name]           | [Reference]      | [Open or Closed] |

An exception cannot waive applicable law, contract, or a non-waivable organization rule. A conditional decision remains blocked until every required condition has verified closure.

## 12. Final Analysis and Decision

Preserve the frozen plan, raw result references, calculation version, exclusions, missing data, incidents, changes, and reviewer notes.

| Gate                   | Threshold    | Observed result | Evidence    | Pass, fail, or unknown | Owner sign-off  |
| ---------------------- | ------------ | --------------- | ----------- | ---------------------- | --------------- |
| Primary outcome        | [Threshold]  | [Result]        | [Reference] | [Status]               | [Name and date] |
| Required control gates | [Thresholds] | [Results]       | [Reference] | [Status]               | [Name and date] |
| Evidence sufficiency   | [Threshold]  | [Result]        | [Reference] | [Status]               | [Name and date] |
| Human workload         | [Threshold]  | [Result]        | [Reference] | [Status]               | [Name and date] |
| Unit and total cost    | [Threshold]  | [Result]        | [Reference] | [Status]               | [Name and date] |
| Stop and recovery test | [Threshold]  | [Result]        | [Reference] | [Status]               | [Name and date] |

Compare every observed result with the baseline and precommitted threshold. Include failed and aborted runs under the approved denominator rule. Explain material uncertainty and any conditions that differed from the intended deployment setting.

Final decision: [Stop | Extend pilot | Narrow | Expand pilot to a named bounded scope | Production candidate]

Decision rationale: [Evidence-based rationale]

Approved next scope: [Exact users, workflow, environment, authority, data, context, tools, volume, dates, and spend]

Required conditions and closure evidence: [References]

Production release still requires: [Risk review, charter, readiness assessment, deployment checklist, approvals, or other gates]

Decision owner: [Name]

Decision date and time: [Timestamp]

Next-scope approval record, required before expansion or further operation:

- Decision: [Approved | Rejected | Returned for changes]
- Authenticated approver identity: [Stable principal ID]
- Authority source and effective role at decision time: [Protected reference]
- Approved plan and result-record version or digest: [Exact versions]
- Approved next scope and exact release: [Stable references]
- Source and target tenant, workspace, and account IDs: [Every applicable boundary]
- Decision date and time: [Timestamp]
- Approval expiry: [Timestamp]
- Conditions, owners, and closure evidence: [Protected references]
- Decision evidence: [Protected reference]

Plan and evidence archive: [Protected reference]

The final pilot decision is an analysis record, not authorization for the next scope. Any failed required gate, prohibited event, unknown material control result, missing required decision, expired decision, or open condition blocks expansion. Do not rewrite the original thresholds after seeing results. An extension that changes the plan, release, or scope requires a new version that records the reason and receives a new approval.
