Choose one job and one owner
Define one audience, one business problem and one person responsible for the pilot. “Improve customer experience” is too broad to evaluate. “Help new customers finish the documented setup and route unresolved issues to support” gives the team something observable.
Set the scope before choosing presentation style. Identify what the employee may answer, which actions it may take and which decisions belong to people.
Record a baseline with clear denominators
Measure the existing process over a representative period. Count eligible customers, completed outcomes and unresolved cases. Write down the inclusion rules so the comparison after launch uses the same population.
For sales, distinguish visitors, enquiries and sales-accepted leads. For support, distinguish conversations, resolutions and reopened issues. For onboarding, distinguish account creation from the first useful product outcome. Do not combine these events into a single success count.
Build a review set before rollout
Prepare representative questions with expected answers and source references. Include ambiguous requests, missing facts and requests outside the employee’s scope. Use ordinary customer language rather than only the exact wording of documentation.
Review actual responses against the source. Record the date, source version, prompt, answer and reviewer decision. Test connected actions separately by inspecting the destination system.
Read related guides
Use a balanced scorecard
Choose one primary business outcome and a few quality checks. Set your own acceptance thresholds based on the risk and current performance. The following are measurement categories, not claimed Aivah results.
- Outcome: accepted lead, resolved issue or activation milestone.
- Quality: correct answer rate in a reviewed sample.
- Customer effort: repeated questions and time to a useful result.
- Handoff: whether the correct person receives sufficient context.
- Reliability: failed actions, duplicate records and connection errors.
- Operating cost: actual usage and the time spent reviewing and maintaining the workflow.
Expand only after the evidence is useful
Compare like-for-like traffic and allow enough completed journeys to make the result meaningful. Small samples, campaign changes and internal tests can make an apparent improvement misleading. A before-and-after comparison alone does not establish that AI caused the change.
Keep a log of improvements and unresolved failures. Expand to another page or workflow when the first job is understood and maintained, rather than adding more responsibilities to an unreliable pilot.
Read related guides

