Standard 12 — Digital Care and Artificial Intelligence
Report content, delivery, and turnaround workflow are addressed in Standard 8. This standard addresses the specific governance of artificial intelligence used in diagnostic result interpretation.
Criteria in this standard
12.2 — A Qualified Professional Genuinely Reviews Every AI-Assisted Result
12.3 — AI-Assisted Tool Performance Is Genuinely Monitored on an Ongoing Basis
12.4 — Laboratory Staff Are Genuinely Consulted and Trained Before an AI Tool Is Introduced
12.5 — Accountability for AI-Assisted Diagnostic Decisions Is Explicitly Defined
12.6 — The Ordering Clinician Is Genuinely Informed When a Result Involved AI-Assisted Interpretation
AI-Assisted Diagnostic Tools Are Genuinely Validated Before Clinical Use
Core
In plain terms: Before an AI tool is used to help interpret real results, the lab has actually tested it on its own patients and conditions — not just trusted the vendor’s general performance numbers.
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
This is marked Core because a diagnostic tool’s published performance figures are genuinely based on the population and conditions it was trained and tested on — which may differ meaningfully from a specific laboratory’s own real patient population, equipment, and sample preparation practices. Local validation is what actually confirms the tool performs as claimed in this specific laboratory’s real conditions, not just in the vendor’s own testing environment.
What good looks like
- A tool is genuinely validated against the lab’s own population.
- Validation is genuinely documented with real figures.
- Re-validation is genuinely triggered by meaningful change.
Common failure modes
- A tool is deployed based solely on the vendor’s published accuracy figures, with no genuine local validation.
Worked example
If you are starting from zero — do this first
- Run a local validation study against this laboratory’s own sample set before any clinical use.
Self-assessment questions
Evidence: Validation study documentation
Evidence: Validation report
Evidence: Re-validation protocol
Common reasons for a PARTIAL answer
- A validation study was conducted but never actually repeated after a significant equipment change.
Implementation plan
| When | What |
|---|---|
| Before any clinical use | Run a local validation study and document results. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| DOCUMENT | Validation study review | Reviews the local validation documentation and figures. |
Supervisor tips
- Ask for the actual local validation data, not just the vendor’s marketing material.
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.
A Qualified Professional Genuinely Reviews Every AI-Assisted Result
Core
In plain terms: A qualified person always actually checks an AI-assisted result before it goes out — the AI never has the final say on its own.
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
This is marked Core because a diagnostic error reaching a clinician and patient genuinely carries real, direct harm potential — AI-assisted interpretation, however well-validated, is a support to professional judgement, not a replacement for it. Genuine, meaningful human review is the actual safeguard against the AI tool’s own real limitations and failure modes.
What good looks like
- Every result is genuinely reviewed before release.
- The reviewer can genuinely override the AI output.
- A real instance shows a genuine override.
Common failure modes
- Review has become a formality — the reviewer habitually approves AI output without genuine, independent scrutiny.
Worked example
If you are starting from zero — do this first
- Introduce a structured review checklist requiring documented, independent assessment.
Self-assessment questions
Evidence: Review records
Evidence: Override capability documentation
Evidence: Override record
Common reasons for a PARTIAL answer
- Review happens but has genuinely become a rubber-stamp rather than independent scrutiny.
Implementation plan
| When | What |
|---|---|
| Week 1-2 | Introduce a structured, documented review checklist. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| DOCUMENT | Review record sample | Reviews a sample of completed reviews for genuine, substantive scrutiny. |
Supervisor tips
- Ask for a real, specific example of an AI interpretation being genuinely overridden.
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.
AI-Assisted Tool Performance Is Genuinely Monitored on an Ongoing Basis
Standard
In plain terms: The lab keeps actually checking whether the AI tool stays accurate over time, not just trusting the one-time validation done before it was first used.
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
An AI tool’s real performance can genuinely drift over time as reagent batches, equipment, or patient population characteristics change — initial validation, however thorough, is a snapshot, not a permanent guarantee. Ongoing monitoring is what actually catches a genuine performance drift before it causes real harm.
What good looks like
- Accuracy is genuinely monitored on an ongoing basis.
- Discrepancies are genuinely tracked.
- A genuine pattern triggers a real response.
Common failure modes
- Initial validation is treated as a one-time check, with no genuine ongoing monitoring afterward.
Worked example
If you are starting from zero — do this first
- Introduce a discrepancy log tracking AI-versus-professional disagreement.
Self-assessment questions
Evidence: Monitoring records
Evidence: Discrepancy log
Evidence: Escalation protocol
Common reasons for a PARTIAL answer
- A discrepancy log exists but is never actually reviewed for genuine patterns.
Implementation plan
| When | What |
|---|---|
| Week 1-2 | Introduce a discrepancy log and a regular review schedule. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| DOCUMENT | Discrepancy log review | Reviews the log and evidence it is genuinely analysed over time. |
Supervisor tips
- Ask when the discrepancy log was last actually reviewed, not just whether it exists.
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.
Laboratory Staff Are Genuinely Consulted and Trained Before an AI Tool Is Introduced
Standard
In plain terms: Before a new AI tool is introduced, staff get a real say and genuinely understand its specific limitations — not just trained to click “approve.”
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
Staff who will genuinely review AI-assisted output need real, substantive understanding of the tool’s specific, known failure modes to actually catch an error — training limited to “how to operate the software” misses this genuine, more important understanding of when and how the tool can actually be wrong.
What good looks like
- Staff are genuinely consulted before introduction.
- Training genuinely covers specific, known limitations.
- Staff can genuinely describe a real limitation.
Common failure modes
- Training covers only how to operate the software, with no genuine understanding of the tool’s actual failure modes.
Worked example
If you are starting from zero — do this first
- Build training content around the tool’s actual, documented failure modes, not just its interface.
Self-assessment questions
Evidence: Consultation record
Evidence: Training materials
Evidence: Staff interview
Common reasons for a PARTIAL answer
- Training happened but genuinely covered only interface operation, not limitations.
Implementation plan
| When | What |
|---|---|
| Before any launch | Build and deliver limitation-focused training. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| ASK | Staff interview | Asks staff to describe a specific limitation of the AI tool. |
Supervisor tips
- Ask a staff member what the AI tool is known to get wrong, specifically.
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.
Accountability for AI-Assisted Diagnostic Decisions Is Explicitly Defined
Core
In plain terms: Everyone knows, in advance, who’s actually responsible if an AI-assisted diagnostic result turns out to be wrong.
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
This is marked Core because a genuine diagnostic error carries real, direct patient harm potential — leaving accountability genuinely unexamined until an actual misdiagnosis occurs means the question gets answered under real pressure, potentially poorly, exactly when clarity matters most. Explicit, documented accountability, established before any incident, is what ensures a genuine, prompt, and appropriate response.
What good looks like
- Accountability is explicitly documented.
- Reviewing professionals genuinely understand their own accountability.
- A defined process exists for reviewing accountability in an error.
Common failure modes
- Accountability has never been genuinely examined until an actual diagnostic error forces the question.
Worked example
If you are starting from zero — do this first
- Document accountability for AI-assisted diagnostic decisions explicitly, before an incident forces the question.
Self-assessment questions
Evidence: Documented accountability policy
Evidence: Staff interview
Evidence: Incident review protocol
Common reasons for a PARTIAL answer
- A policy exists but staff genuinely haven’t been told about it.
Implementation plan
| When | What |
|---|---|
| Week 1-2 | Document accountability explicitly and communicate it to all reviewing staff. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| ASK | Staff interview | Asks a reviewing professional to describe their accountability for an AI-assisted result. |
Supervisor tips
- Ask a reviewing professional directly: “If the AI was wrong and you approved it, who’s responsible?”
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.
The Ordering Clinician Is Genuinely Informed When a Result Involved AI-Assisted Interpretation
Standard
In plain terms: The lab makes sure the doctor who ordered the test actually knows when AI helped interpret the result, so the doctor can tell the patient — since the lab itself usually has no direct relationship with the patient.
| Facility category | Crisis | Transition | Small | Standard |
|---|---|---|---|---|
| Applicability | Adapted | Full | Full | Full |
Why this matters
A laboratory genuinely sits one step removed from the patient — the ordering clinician is the actual point of contact who explains results. A patient’s genuine right to know when AI was involved in their care can only be honoured here if the laboratory first ensures the clinician genuinely knows, since the lab has no direct channel to disclose to the patient itself. This criterion is the laboratory’s own half of an obligation it shares with the ordering clinician.
What good looks like
- The report or an accompanying note genuinely flags AI involvement.
- This disclosure is genuinely clear enough for the clinician to pass along.
- A clinician can genuinely confirm they noticed it.
Common failure modes
- AI involvement is technically logged in an internal system but never genuinely appears on or alongside the report the clinician actually receives.
Worked example
If you are starting from zero — do this first
- Add a standard, brief disclosure line to the report template for any result involving AI-assisted interpretation.
Self-assessment questions
Evidence: Report template
Evidence: Report wording sample
Evidence: Clinician interview
Common reasons for a PARTIAL answer
- A disclosure line exists in the template but clinicians genuinely don’t notice it among other report text.
Implementation plan
| When | What |
|---|---|
| Week 1-2 | Add a standard AI-disclosure line to the report template for any AI-assisted result. |
How the Monitor verifies this
| Method | What | Detail |
|---|---|---|
| ASK | Clinician interview | Asks an ordering clinician whether they noticed AI disclosure on a recent report. |
Supervisor tips
- Ask to see a real, recent report with AI involvement and check the disclosure is genuinely visible, not buried.
Evidence base
ASF training courses on GMJ Academy →
Foundation courses A-00 to A-03 are live. Criterion-specific modules are being developed and will link here when published.