Foundry Academy · AI Workflow Training · Lesson 5 of 6

Evaluation sets and quality metrics

Build a representative evaluation set and use decision-linked metrics to determine whether a workflow is fit for its bounded purpose.

Start the lesson

Your learning work, on this device

No signup, cloud storage, cross-device sync or verified completion. Saving is optional. This browser profile is shared with anyone who can use it; private mode, browser cleanup or storage limits may remove work. Use only fictional or non-sensitive material. Export a copy before relying on this device.

Not saved. Worksheets, answers and practice notes currently last only in this tab.

Practice markers are self-reported, never credentials.

01 · Explanation

Evaluation sets and quality metrics

Objective: Build a representative evaluation set and use decision-linked metrics to determine whether a workflow is fit for its bounded purpose.

An evaluation set should represent ordinary work, difficult cases, rare but harmful failures, missing information, conflicting sources, multilingual input where relevant, and attempts to bypass instructions. Preserve the inputs, authorized sources, expected qualities, critical failures, and scoring rubric under a version. Keep a holdout subset from routine prompt tuning so improvement is not measured only on familiar examples. Use synthetic data unless approved real records are necessary and governed. Evaluation should include groups or conditions that could experience different error rates, while avoiding unsupported demographic inference.

Choose metrics based on the decision. Useful measures may include source-supported claim rate, critical-failure rate, abstention quality, extraction accuracy, reviewer agreement, override rate, processing time, and cost per accepted output. An overall average can conceal unacceptable high-impact errors, so define non-compensable failures and release thresholds in advance. Compare against the current human or non-AI process when feasible; faster is not better if rework or harm increases. Record confidence intervals or sample limitations rather than presenting small tests as certainty. Re-evaluate after material changes and monitor real exceptions because a pre-release set cannot represent every future condition.

Before you begin

  • Confirm F02-F04, F08, and F09; record that no forty independently labeled examples, approved thresholds, adjudicator assignment, test execution, or result set is supplied.
  • STOP. If the forty-case set, independent labels, thresholds, adjudicator assignment, or result set is missing or conflicts with F02–F04/F08/F09, route the evaluation to the authorized evaluation owner and adjudication reviewer; do not claim threshold approval, test execution, or performance results.

Original overview module anchor →

02 · Compare the artifacts

Supported work. Visible uncertainty.

This is a fictional, sanitized AI-governance training case. All records are fabricated, and the exercise provides no legal, privacy, employment, medical, financial, security, or model-performance guarantee.

Redwood Assist governed intake workflow

Redwood Assist is a fictional services company considering an AI workflow to summarize new client intake and draft a routing recommendation. The current form receives about 420 submissions each month. It includes name, business email, phone, company, project description, budget range, requested timing, and an optional attachment. In a sample of 60 fabricated records, seven attachments contain government identifiers, four contain medical details unrelated to the service, nine include third-party personal information, and 16 lack enough context for reliable routing. The proposed model vendor offers a standard account with model-improvement use enabled by default and a 30-day content retention statement. No security, privacy, contract, data-location, deletion, or subprocessors review has been completed. A pilot prompt tells the model to identify serious founders and reject low-quality leads. The output labels applicants high, medium, or low potential, but neither potential nor serious is defined. In a 40-record test, the model routes 31 correctly according to one manager, while a second manager disagrees on 11 of those decisions. Two Spanish-language submissions are summarized with missing budget conditions. The team proposes automatically rejecting low-rated applicants and sending their attachment to a funding partner. There is no consent for partner disclosure, no human-review standard, no protected exception path, and no versioned evaluation set. Leadership wants the automation live next week to save eight hours of staff time. Learners must classify risk, constrain instructions and sources, establish data boundaries, design meaningful human review, build an evaluation set, and control exceptions and changes. The exercise cannot establish vendor suitability, lawful processing, fairness, or permission to automate decisions.

Supported example — reference only

Evaluation ID
EV-01
Risk category
Privacy leakage
Test-input classification
Fabricated attachment class containing a government-identifier pattern; no live identifier.
Independent reference label
Pending - independently label before any run.
Expected route/output
Quarantine for qualified review; no external transfer.
Metric
Privacy-leakage critical failures / privacy-leakage cases
Critical-failure flag
Yes - provisional pending approval
Adjudicator
Pending - authorized privacy adjudicator role not assigned
Result
Pending - no test executed
Status
Draft - label and adjudication pending

A well-handled evidence gap

Evaluation ID
EV-02
Risk category
Unsupported claims
Test-input classification
Missing-context record class based on F04; test specimen not supplied.
Independent reference label
Pending - reference set not supplied.
Expected route/output
Abstain and list missing fields; no invented route or claim.
Metric
Unsupported claims / evaluated outputs
Critical-failure flag
Learner proposal - pending approval
Adjudicator
Pending - independent adjudicator not assigned
Result
Pending - no test executed
Status
Open - specimen, label, and approval not supplied

Flawed approach — do not copy

Marking this forty-row evaluation plan “approved and complete” without the required evidence or reviewer is a flawed submission. Stop evaluation or release when reference labels are not independent, sensitive data is live/unapproved, a critical failure occurs, or adjudication is unresolved.

Repair: Rework the forty-row evaluation plan as an evidence-backed draft, not an approved result. Reserve forty stable EV identifiers and assign balanced proposed test classes before adding examples. Define five separate risk categories: extraction, routing, unsupported claims, privacy leakage, and critical failure. For each row, classify only fabricated or sanitized test input and cite the aggregate fact motivating that class. Check the revision against this requirement: Exactly forty EV rows exist with designed category coverage and pending results. If the required evidence is still absent, keep the decision blocked and identify the missing input or authorized reviewer.

Full case record, ambiguities and all assignments →

03 · Bounded practice

Build the forty-row evaluation plan.

Build a versioned test set with routine, difficult, multilingual, missing-data, and critical-failure cases.

Deliverable: A blank forty-row evaluation-plan template, populated risk-category and metric definitions, and one failure-analysis starter anchored to the supplied aggregate evidence.

Complete a bounded starter and gap analysis using only CB01, F02, F03, F04, F08, F09, and the assignment-scope record below. Populate supported fields, label every unavailable field “not supplied,” and cite the input ID for each material statement. You may design a proposed template, control, question, or decision rule, but must label it as a learner proposal rather than observed case evidence. Do not contact people, access live systems, run tests, sign records, claim approval, or invent names, dates, quotations, transactions, results, or source documents.

Exact supplied inputs for this assignment
  • M05-I01 · F02 — Seven of sixty fabricated attachments contain government identifiers.
  • M05-I02 · F03 — Four contain unrelated medical details and nine contain third-party personal information.
  • M05-I03 · F04 — Sixteen sampled records lack enough context for reliable routing.
  • M05-I04 · F08 — Managers disagree on eleven of thirty-one supposedly correct routing decisions.
  • M05-I05 · F09 — Two Spanish-language summaries omit budget conditions.
  • M05-B01 · CB01 — Use CB01, the full versioned case brief printed once at the start of this packet, as a citable narrative source for details not normalized into F01–F12. Preserve its uncertainty language and do not treat narrative detail as approval, complete operational records, or professional judgment.
  • M05-S01 · F02, F03, F04, F08, F09 — Build a starter version of “A blank forty-row evaluation-plan template, populated risk-category and metric definitions, and one failure-analysis starter anchored to the supplied aggregate evidence.” from the listed case facts. Treat requested structures, controls, questions, calculations, and templates as learner-designed proposals. Where an operational record or result is absent, add a gap entry naming the missing evidence and authorized owner instead of fabricating it.

Operating procedure

  1. Reserve forty stable EV identifiers and assign balanced proposed test classes before adding examples.
  2. Define five separate risk categories: extraction, routing, unsupported claims, privacy leakage, and critical failure.
  3. For each row, classify only fabricated or sanitized test input and cite the aggregate fact motivating that class.
  4. Keep the independent reference label pending until a qualified person labels it without seeing model output.
  5. Specify expected route/output and a category-specific numerator and denominator.
  6. Mark critical-failure logic as provisional and assign an adjudicator role before any run.
  7. Record result as pending because no test is executed; final-QC all forty rows for coverage, independence, evidence boundaries, and version.
Field-by-field guidance
Evaluation ID
Use the stable EV row identifier. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Risk category
Choose a defined risk category; do not infer a sensitive attribute. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Test-input classification
Describe the designed test class, not real personal data. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Independent reference label
Record pending until independently labeled and adjudicated. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Expected route/output
State the approved expected behavior as a proposal pending owner review. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Metric
Name the category-specific measure and its denominator. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Critical-failure flag
Mark provisional yes only for the defined severe harm condition. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Adjudicator
Name the qualified adjudication role; do not invent an individual. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Result
Record pending unless the packet explicitly supplies an observed result. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Status
Use a truthful state such as draft, open—not supplied, review pending, or blocked. Module use: Use the plan to design a forty-case evaluation with independent labels and explicit failure measures, not to invent cases or results.
Forty-row evaluation plan · learning draft
Evaluation IDRisk categoryTest-input classificationIndependent reference labelExpected route/outputMetricCritical-failure flagAdjudicatorResultStatus

Start with 6 rows; the complete workbook specifies 40 stable rows for this artifact. Add rows here or use the full download. No action is saved until you explicitly choose saving above.

Download complete six-module workbook (.md) · Structured case packet (.json)

Keep private client data, unpublished inventions, personal identifiers and credentials out of these public learning tools.

Module 5 · 2-item formative check

Evaluation sets and quality metrics

Choose an answer and request feedback. Read why each option does or does not fit the evidence. Answers stay in this tab unless you choose device-only saving; they are never submitted.

Question 1 of 2 · MODULE 5 · knowledgeWhy must an AI evaluation set include difficult cases and defined critical failures?
Question 2 of 2 · MODULE 5 · scenarioF02 and F03 confirm sensitive-data examples, F04 confirms insufficient context, F08 is conflicting across supposed correct labels, and F09 confirms omitted Spanish budget conditions. What evaluation design may the learner prepare?

Answer either question to review its reasoning.

Inspect the artifact, not just your quiz answers

  • Exactly forty EV rows exist with designed category coverage and pending results.
  • Independent-label and adjudication fields are never backfilled with invented outcomes.
  • Five category-specific measures and provisional thresholds are reviewable.

Stop: Stop evaluation or release when reference labels are not independent, sensitive data is live/unapproved, a critical failure occurs, or adjudication is unresolved.

Go: Proceed to controlled evaluation only after the versioned set, labels, metrics, thresholds, owners, and privacy controls are approved.

Escalate: Escalate label disputes, privacy leakage, unsupported claims, and threshold changes to independent adjudication and the accountable risk owners.

04 · Evidence to keep

Leave with usable work.

Submit the versioned set, rubric, critical-failure definitions, results, limitations, comparison, and release recommendation.

Download your artifact CSV and, if wanted, export the learning-work JSON above. Neither export is a reviewed submission or certificate. Device-only saving is optional; you must press Save my work now after edits.

When all six artifacts are ready, compare the full packet against the track rubric. Qualified human review is still required before real-world decisions.

Technology team discussing an AI-assisted workflow and its controls.
Learn the standard. Practice the work.
Professional reviewing an AI-assisted output on a laptop before approval.
Leave with evidence you can inspect.

Sources, scope and review boundaries

Curriculum 2026.10.08-learning-paths-1. External source dates below are record checks, not continuing guarantees. Verify current requirements before consequential use.

nist-ai-rmf · Official guidance

NIST Artificial Intelligence Risk Management Framework

Primary NIST resource for governing, mapping, measuring, and managing AI risk across the lifecycle.

Open reviewed external source ↗

nist-ai-600-1 · Official guidance

NIST AI 600-1: Generative Artificial Intelligence Profile

NIST's cross-sector profile describing generative-AI risks and actions aligned with AI RMF 1.0.

Open reviewed external source ↗

ws-ai-workflow-operating-standard · Academy internal operating standard

Wealth Synergy AI workflow internal operating standard

Academy-selected task classification, prompt, evidence, human-review, evaluation, exception, and change controls. This is an internal operating standard selected by Foundry Academy; it is not law, accreditation, licensure, or an external-standard requirement.

Version 1.0 · reviewed 2026-09-01 · owner: Foundry Academy curriculum owner

A future Wealth Synergy private professional-development certificate would be issued only after its assessment, capstone, identity, reviewer, retention, access, deletion, appeal, and issuance controls pass quality review. No credential is currently issued. Any future certificate would not be an accredited academic qualification, professional license, or government certification.