Pick a task you already know how to do and can check. The first pilot should produce a draft for review, not send a message, change a record, or publish a page. You will learn more from one complete task than from a long list of possible uses.
The 30-minute run
| Time | What to do | What to keep |
|---|---|---|
| 0–5 minutes | Choose one recurring task and write what a good result must contain. | A one-sentence success rule. |
| 5–10 minutes | Prepare a small input. Remove private information or use a fictional example. | The exact input and permission decision. |
| 10–18 minutes | Ask the tool for one draft with clear limits. | The prompt and first output. |
| 18–25 minutes | Compare the output with the source and your success rule. Correct facts and missing details. | A list of edits and errors. |
| 25–30 minutes | Decide whether to repeat, change, or stop the experiment. Count setup and review time. | A next step, owner, and date. |
Worked example: sort reader questions
Imagine a small garden shop receives five messages each morning. The owner wants a short internal summary, grouped by product question, delivery, or return. The example is fictional so it can be tried without real customer data.
Input: “Is the blue planter in stock?” “My order arrived cracked.” “Can I return an opened bag of soil?” “When will order 123 arrive?” “Do you carry gloves for children?”
Success rule: Include all five messages once, keep the original wording available, flag the cracked item and opened-return question for a person, and never invent inventory or delivery facts.
Prompt:
Group these five fictional customer questions into product, delivery, and return. Make a table with the original question, category, and a short next step. If a question involves damage or a policy exception, mark it Needs review. Do not draft a customer reply, assume stock status, look up an order, or invent a policy.
Review: Check that no question disappeared, was duplicated, or received a made-up answer. If “arrived cracked” is treated as an ordinary delivery question with an automatic reply, the pilot failed the escalation rule. If the summary is accurate, the next test could use a permitted, de-identified set of past questions. Sending replies is a separate decision.
Copyable pilot worksheet
- Task and owner: ______________________________
- Current steps and approximate full-run time: ______________________________
- Good result in one sentence: ______________________________
- Input and sharing permission: ______________________________
- Prompt and tool/account used: ______________________________
- Facts or fields that must be checked: ______________________________
- What the tool may do: Draft / summarize / classify / other: __________
- What requires approval: ______________________________
- Errors, omissions, and correction time: ______________________________
- Total setup + run + review time: ______________________________
- Decision: Repeat / revise / stop. Next review date: __________
Print this page from your browser to use the worksheet on paper. For a fair comparison, include setup, checking, rework, and upkeep along with the tool’s run time. A single successful result is evidence that the task is worth testing again, not proof that the workflow is ready to act alone. Count the whole cost of a new tool explains the comparison in more detail.
Before the next run, read what to share with an AI tool, how to verify an AI answer, and how to test a workflow before it acts.