02 · Compare the artifacts
Supported work. Visible uncertainty.
This is a fictional, sanitized AI-governance training case. All records are fabricated, and the exercise provides no legal, privacy, employment, medical, financial, security, or model-performance guarantee.
Redwood Assist governed intake workflow
Redwood Assist is a fictional services company considering an AI workflow to summarize new client intake and draft a routing recommendation. The current form receives about 420 submissions each month. It includes name, business email, phone, company, project description, budget range, requested timing, and an optional attachment. In a sample of 60 fabricated records, seven attachments contain government identifiers, four contain medical details unrelated to the service, nine include third-party personal information, and 16 lack enough context for reliable routing. The proposed model vendor offers a standard account with model-improvement use enabled by default and a 30-day content retention statement. No security, privacy, contract, data-location, deletion, or subprocessors review has been completed. A pilot prompt tells the model to identify serious founders and reject low-quality leads. The output labels applicants high, medium, or low potential, but neither potential nor serious is defined. In a 40-record test, the model routes 31 correctly according to one manager, while a second manager disagrees on 11 of those decisions. Two Spanish-language submissions are summarized with missing budget conditions. The team proposes automatically rejecting low-rated applicants and sending their attachment to a funding partner. There is no consent for partner disclosure, no human-review standard, no protected exception path, and no versioned evaluation set. Leadership wants the automation live next week to save eight hours of staff time. Learners must classify risk, constrain instructions and sources, establish data boundaries, design meaningful human review, build an evaluation set, and control exceptions and changes. The exercise cannot establish vendor suitability, lawful processing, fairness, or permission to automate decisions.
Supported example — reference only
- Evaluation ID
- EV-01
- Risk category
- Privacy leakage
- Test-input classification
- Fabricated attachment class containing a government-identifier pattern; no live identifier.
- Independent reference label
- Pending - independently label before any run.
- Expected route/output
- Quarantine for qualified review; no external transfer.
- Metric
- Privacy-leakage critical failures / privacy-leakage cases
- Critical-failure flag
- Yes - provisional pending approval
- Adjudicator
- Pending - authorized privacy adjudicator role not assigned
- Result
- Pending - no test executed
- Status
- Draft - label and adjudication pending
A well-handled evidence gap
- Evaluation ID
- EV-02
- Risk category
- Unsupported claims
- Test-input classification
- Missing-context record class based on F04; test specimen not supplied.
- Independent reference label
- Pending - reference set not supplied.
- Expected route/output
- Abstain and list missing fields; no invented route or claim.
- Metric
- Unsupported claims / evaluated outputs
- Critical-failure flag
- Learner proposal - pending approval
- Adjudicator
- Pending - independent adjudicator not assigned
- Result
- Pending - no test executed
- Status
- Open - specimen, label, and approval not supplied
Flawed approach — do not copy
Marking this forty-row evaluation plan “approved and complete” without the required evidence or reviewer is a flawed submission. Stop evaluation or release when reference labels are not independent, sensitive data is live/unapproved, a critical failure occurs, or adjudication is unresolved.
Repair: Rework the forty-row evaluation plan as an evidence-backed draft, not an approved result. Reserve forty stable EV identifiers and assign balanced proposed test classes before adding examples. Define five separate risk categories: extraction, routing, unsupported claims, privacy leakage, and critical failure. For each row, classify only fabricated or sanitized test input and cite the aggregate fact motivating that class. Check the revision against this requirement: Exactly forty EV rows exist with designed category coverage and pending results. If the required evidence is still absent, keep the decision blocked and identify the missing input or authorized reviewer.
Full case record, ambiguities and all assignments →