A marketing team can produce more work with AI and still have little evidence that the business improved. Faster drafts, completed tool calls and enthusiastic feedback answer useful questions, but they do not answer the same question as incremental profit. A good pilot creates enough evidence to decide whether to expand a particular workflow, change its design or stop it. Define that decision before selecting the success metric.

Choose one business decision, not an adoption target

“Use AI in marketing” gives a team no meaningful boundary. A more useful pilot might prepare campaign briefs, classify incoming inquiries or propose recovery messages for review. Each has a different population, quality standard and route to commercial value. Choose one workflow whose beginning and ending can be observed. Name the person who will decide what happens after the pilot.

Write a short decision statement: we will consider expanding this workflow if it meets an agreed quality threshold, improves a specified operational measure and stays within defined cost and risk limits. Set the thresholds before seeing the results. They should come from your economics and service expectations; there is no universal percentage that makes an AI pilot successful for every business.

Separate three layers of evidence

Our proposed evaluation model separates task quality, operating economics and commercial outcomes. Task quality asks whether the work was correct and useful. Operating economics asks what it cost to get acceptable work completed. Commercial outcomes ask whether customer behavior or business results changed. Keep separate scorecards so that success on one layer does not silently stand in for success on another.

Anthropic's January 2026 evaluation guide distinguishes test tasks, repeated trials and grading logic for agents. It is a useful technical reference for assessing behavior, not evidence that an agent will improve marketing revenue. Anthropic's evaluation guide. Our business scorecard extends the question to reviewer effort, destination outcomes and the decision facing the marketing owner.

Build a baseline that can survive comparison

Record how the current process handles a comparable set of work. Include preparation, review, corrections, waiting and exceptions. If the existing process has no reliable time records, disclose that weakness and collect a baseline rather than manufacturing a precise savings claim. Use the same definition of “completed” for both approaches. A draft waiting for approval is not equivalent to an approved asset.

Also record who did the work and which inputs were available. An experienced employee with a complete brief is not a fair comparison for an AI workflow receiving incomplete requests. Separate straightforward tasks from unusual cases instead of allowing a changing mix to dominate the average. Retain the denominator: report accepted outputs out of all eligible attempts, including failures and abandoned cases.

Design a quality rubric before the demonstration

For a campaign brief, a useful rubric might inspect fidelity to the offer, audience relevance, support for claims, completeness of required fields and usability by the next team. Define what makes a defect serious enough to reject the output. Reviewers should apply the same criteria to both baseline and pilot work. Where practical, conceal the production method during assessment to reduce expectation effects.

Use examples from the actual workflow with appropriate data permissions. Include missing information, contradictory instructions and requests outside the allowed scope. Repeat selected cases because a single good output may not represent stable behavior. Review disagreements between evaluators: they may reveal an unclear rubric or a disputed business policy rather than a model failure. Resolve those ambiguities before treating the score as precise.

Count the cost of accepted work

Measure cost per accepted output, including model or software charges, preparation, human review, rework and relevant integration effort. Keep one-time implementation costs visible separately from recurring operating costs. Explain how shared costs were allocated. An apparently inexpensive output becomes less attractive if senior staff spend substantial time correcting it or investigating incomplete actions.

Consider an illustrative calculation, not a Nextriad customer result. A team spends 600 monetary units on a pilot and accepts 40 outputs. Its measured cost is 15 units per accepted output within that cost boundary. Dividing by 100 attempted outputs would produce 6 units, but would answer a different question. Neither number establishes profitability; the accepted work still needs to create value for the business.

Test commercial impact only when the design supports it

If the workflow affects a customer-facing experience, consider a controlled comparison with a suitable assignment unit and observation window. A specialist should assess sample requirements, overlap between groups and whether the test can detect a useful difference. Keep simultaneous offer, budget and targeting changes visible. If they cannot be separated, describe the result as observational rather than claiming the AI caused the movement.

Use the full relevant outcome window. A faster response can coexist with more unqualified meetings or later cancellations. Track the downstream measure that matters to the workflow, such as qualified opportunities or contribution after returns, where measurement permits. Small samples and delayed outcomes may leave the commercial question unresolved. That is an acceptable conclusion when operational evidence still supports a bounded next step.

Make expansion conditional and reversible

Before the pilot begins, agree what leads to expansion, redesign or suspension. Expand a demonstrated workflow within a defined scope, then check whether new languages, customer types or workloads change performance. Freeze material configuration details in the record so that later results can be linked to the version actually tested. Changes in tools, prompts or policies may require a new evaluation.

NIST's AI Risk Management Framework offers a voluntary reference for managing and evaluating AI risks; this scorecard does not establish conformity with it. NIST AI RMF. For a Nextriad ARS assessment, bring one workflow and its baseline. The first useful conversation is about the evidence needed to make an investment decision, including which promised benefits your pilot will not yet be able to prove.

  • State the workflow, population, owner and decision date.
  • Record baseline quality, total effort and exception handling.
  • Set acceptance criteria and count every eligible attempt.
  • Separate recurring costs from implementation investment.
  • Document commercial attribution limits and delayed outcomes.
  • Define the next scope and the conditions for stopping it.

Sources and review date

Sources reviewed: 2026-09-15.

Put this guide to work

Blank working template; complete with your own evidence. Do not put customer personal data in shared copies.

Download CSV template

Published by Nextriad. Editorial standards and corrections