Which tasks should AI complete?
A customer wants to compare a product, request a quote or track a delivery. Task selection defines which of these journeys an AI agent should test, before measurement starts. It specifies the goal, access and permitted endpoint. Multiple tasks reveal different barriers; a single test remains a limited observation. This shows where Hyperize can improve the customer journey.
Fix the task before the test. Read the result within its tested scope.
The starting point
Start with the customer need.
One failed journey does not describe an entire brand. The task must reflect a real customer need.
The methodology separates selection from measurement: define the goal and limits, then test the journey.
A single website check covers one slice. A broad brand comparison needs multiple comparable tasks.
Six rules for fair tests
Rules for the test plan, not proof that every historical measurement meets every criterion [S1].
01
Test a real customer need.
02
Test access within the brand’s responsibility.
03
Use multiple tasks for broad comparisons.
04
Keep difficulty comparable within a sector.
05
Separate intermediary routing from technical barriers.
06
Freeze and publish task descriptions before the comparison wave.
Avoid five errors
These errors distort a comparison [S1].
Test the appropriate endpoint
Direct surface
The journey at the brand
What can the agent complete on the brand’s own pages and systems?
Intermediaries
The journey via third parties
Does the task require an intermediary, or could the agent have stayed with the brand?
Intermediaries provide context, not an extra factor in the published score formula [S2].
Four task-selection examples
Planning examples, not new brand measurements. Verify the goal and available access for each task.
Insurance
Allianz
Access route · Quote and advice
A suitable inquiry can be the endpoint.
Suitable task
Check coverage; prepare an inquiry.
An unfair requirement
Count an intermediary result as website failure.
Controls
FM1 · FM3
Automotive
Mercedes-Benz
Access route · Configuration and contact
Define the endpoint for each offer.
Suitable task
Check configuration; find a test drive.
An unfair requirement
Require online purchase for every journey.
Controls
FM2 · FM3
Logistics
DHL
Access route · Shipping and tracking
Specify the parcel task and access.
Suitable task
Calculate a price; track a shipment.
An unfair requirement
Treat a price lookup as equivalent to logged-in checkout.
Controls
FM2
Healthcare
Bayer
Access route · Information and supply
Separate product information from purchase.
Suitable task
Find documentation and a supply route.
An unfair requirement
Assume direct manufacturer checkout.
Controls
FM1 · FM3
Comparability
Freeze tasks. Explain changes.
For comparison waves: freeze and publish task descriptions before testing. Retests add to the history [S1].
A changed task gets a new version. Compare trends only under comparable conditions.
Findings can be challenged with evidence. Corrections need a traceable change record.
Scope and limits
What the test establishes.
Rules and score weights are public. Internal prompts, scoring details and raw logs stay within the project. A planned stop before purchase or submission is not a technical failure. Examples do not establish current brand performance; rankings and citations are not guaranteed.
Evidence and provenance
S1
Internal
Task selection rules
Hyperize · Methodology · May 2026
Internal methodology, not publicly retrievable.
Supports: Six rules, five failure modes and versioned task descriptions.
S2
Methodology
Agent Success Score
Hyperize · October 2026
https://www.hyperize.ai/en/methodology/agent-success-score
Supports: Public 20/70/10 weights and measurement limits.
S3
Evidence
DHL: historical task example
Hyperize · DAX 40 Index · Q2 2026
https://www.hyperize.ai/en/dax40-index/brands/dhl
Supports: One parcel task; checkout behind login untested. Not a complete task mix.
S4
Evidence
DAX 40 Index
Hyperize · Published measurement periods
https://www.hyperize.ai/en/dax40-index
Supports: Results within their stated scope and limits.
Internal methodology is not publicly retrievable. Public examples document their own test scope.
From testing to implementation
Measurement
Agent Success Score
What the score summarizes and where its limits lie.
Understand the scoreMy next step
My customer task
Agree the goal, access and scope for your project.
Discuss my customer task