The short answer
What matters most
Plan an AI pilot around one task, a representative evaluation set, and an explicit decision about human review. Compare quality, review effort, turnaround time, and operating cost with the current process. The pilot should determine whether to stop, revise, or proceed to a limited rollout—not merely demonstrate that the model can produce a plausible answer.
- Test the difficult exceptions alongside normal inputs.
- Separate a helpful draft from an authorized automated action.
- Define acceptance and stopping criteria before the demonstration.
Compare before you commit
Decision at a glance
| Pilot question | Evidence to collect | Decision it supports |
|---|---|---|
| Does it produce usable output? | Field-level errors, missing facts, and unsupported statements | Whether output quality is sufficient for the selected task |
| Does it reduce work? | Review and correction effort compared with the current process | Whether the overall workflow has a practical benefit |
| Can failures be contained? | Escalation paths, permissions, and recovery examples | Where a person must remain in control |
| Can it operate reliably? | Latency, usage cost, integration errors, and monitoring needs | Whether to progress to a bounded rollout |
Is this workflow ready for a pilot?
Start when someone can provide representative examples, judge the output, and describe the action that follows. If inputs are inaccessible or the business rule is unsettled, resolve that first. If structured data and explicit rules cover the task, compare conventional automation before adding a model.
Discuss your automation use caseChoose one bounded task
A useful pilot starts with one job that people can describe and judge: extracting fields from a document, drafting a summary, routing an enquiry, or finding an answer in an approved knowledge base. Record the input, expected output, user, and action that follows. Avoid making the first objective a broad transformation of an entire department.
Check whether ordinary workflow automation can solve the problem. If every decision follows explicit rules and the inputs are structured, a conventional application may be simpler. AI becomes more relevant when language or varied documents are central to the task. The pilot should test that choice, not assume it.
Build a representative evaluation set
Collect examples that reflect normal work and difficult exceptions. Include incomplete inputs, conflicting information, unusual formats, and cases where the correct response is to ask for help. Obtain permission to use the data and remove information that is unnecessary for the test.
Have someone who understands the task describe what a good result looks like. For extraction, check fields individually. For summaries, check missing facts and unsupported statements. For question answering, check whether the response is supported by the permitted source material. Keep a separate set of examples for evaluation so improvements are not judged only on familiar samples.
Define review and access boundaries
Decide what the system can read, what it can change, and who can see its output. A pilot that drafts a response has a different risk from one that sends it automatically. Start with human review where a mistake could materially affect a customer, financial record, or clinical workflow.
In Salt’s clinical intake project, recorded responses, timestamped transcripts, and structured summaries support clinician review. That is an example of AI assisting documentation within a broader application. It does not establish that an unrelated workflow can safely run without supervision.
Measure usefulness as well as accuracy
Choose acceptance criteria before the demonstration. Include error types, reviewer effort, turnaround time, and operating cost. A fluent result may still require enough correction to erase the benefit. Compare the proposed workflow with the current process using similar examples.
Record failures instead of hiding them in an average. Missing a critical field may matter more than several formatting mistakes. Define which errors stop a release, which route the task to a person, and which are acceptable within a clearly communicated limitation.
Plan the pilot deliverables
The scope should include the workflow, approved data sources, evaluation examples, results, known failure modes, and a recommendation for the next step. Define access to the prototype and ownership of any custom code or configuration. Budget for integration and review work as well as model usage.
Agree on a decision point: stop, revise the task, or proceed to a limited rollout. A successful demonstration is not the same as production readiness. Wider use may require stronger permissions, monitoring, incident handling, and repeated evaluation when documents, prompts, or models change.
Illustrative example · not a client result
Extracting fields from supplier documents
An operations team receives documents in different formats and manually enters supplier details.
- Choose the fields that matter and define which omissions must stop processing.
- Prepare approved examples including blurry files, missing pages, and conflicting values.
- Route uncertain or invalid outputs to a reviewer and measure correction effort.
The decision
Proceed only if the end-to-end process is useful and failures are handled acceptably. This is an illustrative workflow, not a claim of a delivered client outcome.
Planning with a US team
For a US buyer, identify who can approve data use and which vendor-processing and retention conditions apply to the project. Include internal security and domain stakeholders before sharing production data. The appropriate controls depend on the workflow and organization.
Use this in your next discussion
Your planning checklist
Work through these points with your stakeholders or shortlisted partner. Record the owner of any unanswered question.
- Define the task, users, input, output, and next action.
- Obtain approved data and separate evaluation examples.
- Specify errors that require review or stop processing.
- Measure reviewer effort as well as output quality.
- Document results, limitations, access, and the rollout decision.
Common questions
Does every automation project need AI?
No. Structured inputs and explicit rules may be better handled by conventional workflow automation. Test whether AI is needed for interpretation of language, documents, or other variable inputs before adding it.
When can we remove human review?
Only after assessing the consequences of error, demonstrated performance, permissions, monitoring, and recovery for the specific task. A successful pilot alone is not enough to justify autonomous action across a broader workflow.
What happens after a successful pilot?
Plan a limited rollout with defined users, monitoring, support, and evaluation. Recheck performance when source data, prompts, models, or workflows change. Production integrations and operating controls may require additional work.
What should I share with Salt to start this discussion?
A short description is enough: one task and the people doing it today; input formats and a description of common exceptions; what a useful result would change for the team. You can leave confidential documents out of the initial enquiry. We will clarify the requirements and agree whether a project, team, or assessment is the appropriate next step.
Explore Salt’s work
Relevant project experience
The clinical intake project shows AI-assisted documentation for clinician review. KnowledgeBot addresses access to medical protocols. These are distinct workflows that help illustrate why evaluation and human oversight must match the task; they are not evidence of general-purpose autonomous AI.
AI-Assisted Intake Documentation for US Behavioral Health
Salt built a web application for a healthcare technology team serving addiction treatment facilities in the US. It connects recorded intake assessments, transcripts, and structured documentation for clinician review.
Read the case study ↗Healthcare & behavioral healthKnowledgeBot: AI Medical Protocol Assistant for Nurses & Clinicians
Salt’s KnowledgeBot project for Addiction Doctors focused on a custom AI agent that helps nurses and clinicians access medical protocols.
Read the case study ↗Further reading
A voluntary framework for considering risk across AI design, development, use, and evaluation. Use it as additional planning guidance; it does not certify a pilot.
Your next step
Have a repetitive workflow worth testing?
Salt builds AI-assisted applications and workflow automation. Describe the task, the inputs, and what people do with the result. We can discuss whether a rules-based workflow or an AI pilot is the more appropriate starting point.
A useful starting point
- One task and the people doing it today
- Input formats and a description of common exceptions
- What a useful result would change for the team
A few sentences are enough. You do not need a finished specification or confidential documents to start.
Discuss your automation use caseRead client reviews