Chevron left
RESEARCH

Human and AI-assisted decision comparison: a prospective worksheet

An illustrative protocol for a future comparison. Define shared inputs, resource differences, review criteria and observable records before testing human and AI-assisted workflows.
Illustrative brain split between a human form and electronic circuitry; not a measured human-versus-AI result.

Reviewed and updated October 8, 2026.

Use this worksheet to plan a comparison of human work and AI-assisted work on a defined decision. It is an illustrative protocol for a future exercise. It reports no completed participant study, measured advantage or simulation result.

Specify the decision and comparison

Choose a task whose outputs can be reviewed, such as comparing supplier options against documented requirements. Define the decision owner, available evidence, constraints and consequences of a wrong recommendation. A test of recommendations does not authorize the system to place orders or make consequential decisions.

Decide whether you are comparing approaches under equal resources or comparing real operating workflows. For a direct comparison, give participants the same source packet, objectives, permitted tools and time limits. If human consultation, retrieval, tools or computing differ, record those differences and interpret the result as a comparison of those workflows, not a measure of human or model intelligence.

Prospective scenario worksheet

Copy this table into your evaluation plan and complete it before running the exercise.

FieldWhat to specify
Question and scopeDecision to support, output required and actions excluded
Source packetVersioned inputs, missing information, reference facts and permitted external sources
Human conditionParticipant selection, task familiarity, training, allowed consultation and tools
AI-assisted conditionExact model and system versions, prompts, retrieval, tools, settings and human review
ResourcesTime, access, tools, computing and any deliberate differences between conditions
Review methodPre-agreed rubric, independent reviewers, rejected errors and disagreement handling
Run recordCondition order, repetitions, observable actions, outputs, timings, errors and corrections
InterpretationWhat the exercise can answer and what remains outside its scope

Define review criteria before seeing results

Agree the relevant factual requirements and unacceptable errors. Review whether the output uses the supplied evidence, recognizes missing information, follows constraints, compares feasible alternatives and escalates when the evidence is insufficient. Define subjective criteria, such as clarity, with examples so reviewers can apply them consistently.

Keep acceptance cases separate from examples used to tune prompts or train participants. Where practical, hide the condition label from reviewers. Use more than one reviewer for judgment-based criteria and retain disagreements rather than presenting a single score as objective truth.

Record what can be observed

Retain submitted outputs, citations, tool actions, revisions and escalations. A model-generated explanation is not a trace of its internal reasoning. A self-reported confidence score is not a measured probability of being correct. Count generated alternatives only after checking that they are distinct, feasible and relevant.

Measure elapsed time over the same task boundary, including retrieval, tool waits, human review, correction and retries where they are part of the workflow. For streamed output, separate time to first token from time to a usable completed recommendation.

Choose repetition and interpretation deliberately

Choose participants and repetitions to fit the range of cases and the consequence of errors. If the exercise is exploratory or too small to support a broad comparison, say so. Report individual failures and conditions as well as any aggregate. Repeated model runs or simulated scenarios are not independent real-world outcomes.

An exercise can help choose the next evaluation or a division of work. It does not by itself establish general superiority, freedom from bias, business return or safety in deployment. A consequential rollout requires its own acceptance and operating evidence.

Related evidence and next step

The human and AI decision framework outlines the review questions. The voice-grading field report describes a documented grading error and the limits of its evidence. The retrieval proof report shows why retrieval and answer quality need separate checks.

For a scoped evaluation of a real workflow, see custom AI development. Request a project review with the scenario, source packet and acceptance decision the comparison must inform.

Read Our Research