# AI-Assisted GTM Experimentation: A Human-Reviewed Workflow: editable worksheet

All worked records are illustrative. Replace assumptions with your own reviewed inputs.

## Separate preparation from experimentation

AI can draft alternatives, check links and summarise a journal offline. Those are useful preparation tasks, not completed experiments on buyer behaviour. Keep one baseline and one proposed change for the first test.

## A worked example

Illustrative scenario: a B2B team wants to test whether a clearer explanation of an implementation deliverable improves qualified enquiries from a service page.

| Step | Recorded decision |
| --- | --- |
| Baseline | Existing service-page introduction; record its version |
| Hypothesis | Visitors may need a clearer description of what they receive |
| Change | One paragraph; retain the offer, form and traffic sources |
| Primary metric | Qualified enquiries per eligible visitor, using agreed CRM criteria |
| Diagnostic metric | Enquiry submission rate; not a substitute for qualification |
| Guardrails | Broken forms, misleading claims, page errors and accessibility regressions |
| Approval | Named page owner reviews the proposed text and release |
| Outcome | Keep, revert or inconclusive; record the evidence and limitations |

Assign comparable eligible visitors to variants and keep assignment consistent. Where randomisation is unavailable, state the limits of a before/after comparison, including traffic mix and other changes. Do not call the result causal just because the rate moved.

## Agree how to decide before launching

Use the baseline rate, smallest worthwhile effect, available volume and chosen evaluation method to plan the test. There is no universal minimum number of sends or visitors. Record the observation window and allow time for qualification to arrive in the CRM.

Do not promote a variant automatically because it is ahead at a convenient checkpoint. Repeatedly looking for a winner changes the evaluation problem; use a method designed for the agreed checking schedule or review at the pre-agreed endpoint.

For illustration, 3 qualified enquiries from 100 visitors versus 4 from 100 is a small observed difference, not a demonstrated improvement by itself. A journal can record “inconclusive” and retain the baseline. If volume is too low, use interviews and usability reviews to improve the hypothesis rather than manufacturing certainty.

## Define permissions and stop conditions

| Action | Proposed permission |
| --- | --- |
| Draft a variant or summarise observations | AI may prepare it for review |
| Change public copy or release a variant | Named human approves |
| Change budgets, prices or audience eligibility | Separate explicit approval |
| Pause on a broken form or guardrail breach | Pre-authorised stop; alert owner |
| Select a winner or expand to another channel | Human reviews evidence and limitations |

Keep a versioned baseline and a tested restoration path. Decide who can pause the experiment, how they are alerted and how affected observations are excluded before launch.

## Keep the journal useful

Record hypothesis, audience, variant, primary metric, guardrails, evaluation plan, owner, approval, dates, counts, uncertainty and next action. Preserve inconclusive and negative results. Transfer a finding to another channel as a new hypothesis, not as a guaranteed improvement.

## Your working record

| Field | Your input |
| --- | --- |
| Owner | |
| Data sources and dates | |
| Objective | |
| Assumptions to validate | |
| Exceptions | |
| Review date | |
| Decision and supporting evidence | |
| Next action | |
