M.L. Sebastian is now br8n.

Workflow test

Test the handoff before automation

Use this fictional meeting note:

Alex agreed to collect the event photos. No due date was agreed. We discussed a possible homepage refresh; nobody took ownership and no decision to proceed was made.

The task is to turn this note into an action list.

The note contains one agreed action with a missing date. It also contains a possible project that nobody approved. If the answer gives Alex a Friday deadline or assigns the homepage refresh to someone, the handoff added facts that were never in the meeting.

Before you automate a recurring handoff, test whether the system can preserve the difference between a fact, a commitment, and an unanswered question.

Choose one handoff

Start smaller than “automate our meetings.” Choose one result another person has to use, such as:

  • turning a meeting note into an action list
  • turning a customer enquiry into a service brief
  • turning a field conversation into a CRM follow-up
  • turning an approved outline into a first draft

Pick one source and one output. Keep the first test separate from production systems. Use harmless or sanitized information, and do not give the test access to send messages, create tasks, or update records.

The first goal is evidence. You want to see what the handoff preserves, what it changes, and where it guesses.

Write the expected result first

Do this before prompting the model.

If you wait until after the output appears, a fluent answer can change your standard without you noticing. You start grading whether the result sounds reasonable instead of whether it carried the source forward correctly.

Use the companion Workflow Test Card in your browser, or copy the plain worksheet. The same fields also work on paper:

FieldWhat to record
TaskThe one job the AI should perform
Permitted sourceThe exact note, document, or safe example the task may use
Output requirementsThe fields, labels, and source references another person needs
Do not inventOwners, dates, approvals, decisions, numbers, or missing context
Expected resultThe human answer key written before the test
Tool/model and dateWhat ran the test and when it happened
ObservationsThe actual answer, differences, and anything this example did not test
AssessmentNot tested, matches this example, needs revision, or unable to judge
Next actionThe specific rule to change, question to answer, or second example to try

The test needs an answer key before it needs an automation platform.

Build the answer key

For the meeting note above, the desired handoff is an action list that separates commitments from discussion and identifies the questions still blocking a usable next step.

The expected result is:

ItemStatusOwnerDue dateNext question
Collect event photosAgreed actionAlexNot statedAsk Alex for a date
Homepage refreshDiscussed onlyNot assignedNot statedAsk the group whether to proceed

Nothing in the source authorizes the homepage work. Assigning it to a designer would be an invention. Giving Alex a Friday deadline would be another.

The system should preserve those gaps because they are part of the truth of the handoff.

Give the model a boundary it can follow

The first instruction can stay plain:

Create a handoff from the source note below.

Label each item as Agreed action, Discussed only, or Question.
Include an owner or due date only when the source states one.
Write "not stated" when required information is missing.
Never convert a discussion, option, or suggestion into a commitment.
At the end, list the questions needed to complete the handoff.

SOURCE NOTE
[paste the note]

Run it in a separate test. Keep the original source, the instruction, the expected result, and the actual output together. That gives the reviewer a record they can inspect.

Compare the source and answer key

Read the result against the test card.

For this example, fail the test if the output:

  • starts the homepage project
  • assigns an owner who was never named
  • invents a due date
  • drops the question needed to resolve missing information
  • merges the two items into one vague follow-up

The wording can vary. The commitments cannot.

When the result fails, find the smallest rule that would have prevented the error. Add that rule, then run the same source again. Do not quietly repair the output and call the workflow finished. The point is to learn whether the handoff can produce the right boundary repeatedly.

Test the ugly middle

A clean example is a starting point. Real work contains incomplete notes, people with similar names, tentative language, and decisions that depend on another document.

Repeat the test with representative safe examples. Include a missing owner, a missing date, and at least one idea that was discussed but never approved. If the source refers to another document, check whether the system cites that dependency or makes up the missing context.

Record what each test did not cover. A passing meeting-note example does not establish that the workflow can read every document, understand a client account, or update a task system safely. Those are separate checks.

Add automation after the handoff earns it

Only then consider triggers, schedules, task creation, or CRM updates.

Run it in stages:

  1. Produce a draft handoff with no external action.
  2. Have a person compare it with the source and expected result.
  3. Record corrections as rules or exceptions.
  4. Test the same boundary again.
  5. Expand authority only when the evidence supports the next step.

What this test proves

A passing Workflow Test Card shows that one handoff handled one defined example according to written requirements. It can also expose a missing rule or a decision the team never documented.

It does not prove the entire workflow is safe, that another AI provider will behave the same way, or that people will use the finished system. It does not replace access controls, source review, or human judgment.

That narrow result is enough to make a better next decision.

The companion Workflow Test Card keeps the task brief and human review in one record. Its copy button sends only the brief to your clipboard, leaving the expected answer out. The full Markdown download includes the brief, answer key, observations, and assessment so the evidence stays together. The page does not call an AI service or grade its own output. The person testing the handoff still decides whether the result matches the source.

Open the card and load the fictional example. Copy the brief into an AI tool you are allowed to use. Then check three things: Alex owns the photo action, the date stays unstated, and the homepage refresh remains a discussion with no assigned owner. Record the actual answer, download the full card, and try a second safe note before the handoff gets permission to create anything.