The easiest AI demonstration is rarely the best place to start.
Generating a plausible answer from a sample tender can show that a model is capable. It does not show that the workflow solves an important bid problem, works with your evidence or can be trusted when the input is incomplete.
At the other extreme, trying to transform the whole pursuit at once creates too many unknowns. If the pilot fails, you may not know whether the cause was the process, data, model, integration, controls or adoption.
The best first use case is valuable enough to matter, bounded enough to evaluate and safe enough to stop.
Use the process map and lifecycle portfolio from the first two guides. Start with the real points where work waits, repeats, loses context or reaches a weak decision. Then score the candidates before selecting technology.
Begin with a problem statement
Describe the use case without mentioning AI. For example:
Bid reviewers spend too much time checking whether responses address every tender requirement and whether important claims are supported by approved evidence.
That statement identifies a user, pressure and consequence. It does not yet assume the solution.
Now add the decision the workflow must support. In this example, the output may support a proposal-readiness review. The accountable reviewer still decides what must be corrected and whether the response is ready to advance.
If the problem cannot be described clearly, the use case is not ready to score.
Ask Who experiences the problem, when does it occur, and which later decision is weakened by it?
Score five conditions
Use Strong, Conditional or Weak for each condition. The purpose is disciplined discussion, not mathematical precision.
1. Business consequence
Would improvement matter to opportunity quality, capacity, risk, margin protection or delivery confidence? A high-volume nuisance may save time but still be strategically minor. A consequential task can justify investment, provided it is controllable.
2. Recurrence
Does the task happen often enough to learn from? Repetition provides cases for testing and a reason to standardise the workflow. A rare but critical decision may still matter, but it is usually a poor first pilot because evidence develops slowly.
3. Context readiness
Are the inputs available, current, authorised and owned? A promising use case becomes weak when its evidence lives in personal folders, source precedence is unclear or the system cannot distinguish approved material from an obsolete answer.
4. Evaluability
Can a knowledgeable person judge the output against a reliable source or quality standard? Requirements extraction is comparatively evaluable because a reviewer can compare the result with the tender. “Improve the win strategy” is harder because the quality bar is more contextual and the outcome arrives later.
5. Human control and recovery
Can the workflow run without hiding accountability? Name the decision owner, stopping condition, exception route and recovery method. If a material error would remain invisible until after submission, the use case is not a responsible first test.
| Condition | Strong | Conditional | Weak |
|---|---|---|---|
| Business consequence | Changes a meaningful decision, risk or capacity constraint | Useful but indirect benefit | Easy to demonstrate but matters little |
| Recurrence | Enough cases to test and refine | Limited or highly varied cases | Too rare for a useful first learning cycle |
| Context readiness | Sources are approved, current and accessible | Gaps can be repaired within the pilot | Critical context is missing or unreliable |
| Evaluability | Output can be checked against ground truth | Expert judgement can assess it | Success is subjective or visible much later |
| Human control | Owner, gate, stop and recovery are explicit | Controls can be designed before testing | Accountability or recovery is unclear |
Ask Which weak condition would make a convincing demonstration unsafe or irrelevant in real work?
Compare AI with the simpler alternative
Before selecting a candidate, ask whether a clearer owner, earlier meeting, fixed rule, checklist, better search or conventional workflow automation would solve the problem more reliably.
Use AI when language, ambiguity or variable context makes it useful. Decide whether the workflow needs a bounded assistant or an agent with tools and state. This comparison protects the business from solving an operating problem with unnecessary technical complexity.
Three worked examples
These examples illustrate the method. They are not client outcomes or universal recommendations.
Requirements and compliance baseline
Problem: Important requirements are distributed across a large tender pack and later reviews spend time finding omissions.
Assessment: Strong recurrence, context readiness and evaluability when the tender documents are controlled. Human control is clear if a bid professional approves the requirements before they populate the plan.
Potential first test: Extract requirements with source references, compare them with a human-produced baseline and stop before any requirement is treated as approved.
Evidence retrieval for response planning
Problem: Teams repeatedly search for relevant, current and authorised evidence, then reviewers struggle to verify what has been reused.
Assessment: High potential value and recurrence, but context readiness is often conditional. Ownership, permission, currency and source precedence must be defined first.
Potential first test: Retrieve a small evidence pack from one approved collection, with citations, and have an expert score relevance, currency and unsupported matches.
Autonomous Go or No-Go decision
Problem: Qualification is inconsistent and senior time is scarce.
Assessment: Very high business consequence but weak as a first AI use case. Relationship knowledge, strategic fit, capacity, risk and commercial judgement may be incomplete or contested. The decision should remain human.
Potential first test: Prepare a traceable qualification pack and compare whether it makes the human discussion more complete. Do not delegate the pursuit decision.
Write the pilot in one sentence
When [named user] reaches [workflow trigger], the workflow will [bounded AI contribution] using [authorised context] to improve [observable result]. It will stop when [failure or threshold], and [named person] will decide [human gate].
If the sentence is vague, the pilot is still too broad.
Build the smallest credible test
- Baseline. Record the current elapsed time, review effort, quality issues, exceptions and relevant rework.
- Authorised context. Freeze the sources, versions, permissions and rules used by the test.
- Representative cases. Include ordinary work, difficult edge cases, conflicting evidence and poor inputs.
- Shadow mode. Run beside the existing process and do not bypass a formal gate.
- Evaluation. Record omissions, unsupported claims, corrections and reviewer effort against ground truth or expert judgement.
- Failure response. Define when the workflow stops, who investigates and how the team recovers.
- Scale decision. Decide whether to stop, repair, refine, repeat or connect another workflow.
Ask What result would justify the next test, and what finding would make us stop?
Connect the pilot to the practice
An isolated pilot can be technically successful and organisationally useless. Before building, identify what will receive the output and what could reuse the learning.
A reviewed requirements baseline might feed the bid plan, work packages and later assurance. Corrections might improve the evaluation set and source rules. Exceptions might reveal that the process needs repair before a second agent is added.
This does not mean building the whole orchestration layer now. It means designing the pilot so that a useful result can become part of the target workflow rather than another tool the bid team must manage.
The first use case is an investment decision
For the Champion, the pilot must improve a real part of the working day without hiding new review effort. For the Sponsor, it must be consequential enough to justify investment and small enough to generate credible evidence.
That is why selection happens after the process map and lifecycle portfolio. The first use case should prove that one governed workflow can improve, and show what must change before the next connection is earned.
Choose one credible first pilot.
Bring the process map and your shortlist. I can help you test business value, context readiness, evaluability and human control before technology is selected.
Discuss your first use case →Evidence and limitations
This scorecard is an Ignis design method informed by public workflow and agent guidance, bid-practice research and internal practitioner learning. It has not yet been validated as a predictor of client outcomes. Worked examples are illustrative.



