Back to systems
CASE FILE / AF-01REVIEWED EVIDENCEACTIVE SYSTEM / 2025 - Present

AI systems engineering / Human control layer

Astrocade AI QA Calibration Tool

Designed and operated a human-in-the-loop calibration system for Astrocade's AI-powered UGC review pipeline. It samples completed QA reviews, records explicit ground truth and role-scoped diagnostics, and compares QA, final review, and creator-feedback quality through immutable calibration signals.

Operating proof
human-in-the-loop AI systems for creator-facing moderation
Engagement
Astrocade AI
Evidence set
7 reviewed artifacts
Capability coverage
4 documented areas
Human-in-the-Loop AI OperationsQA Annotation & CalibrationModeration Workflow DesignOperational Guardrails

Primary artifact

01 / 07
artifact_viewer.sh

Calibration metrics dashboard

Read-only calibration surface with sampling windows, scoped QA and final-review metrics, agreement, comment quality, and legacy-data exclusions.

01 / Context

Situation

Overview

Astrocade needed fast publishing decisions for an AI-powered UGC platform without treating QA or auto-review as final authority. Quality decisions needed a repeatable handoff from sampled QA evidence to final review and creator-facing feedback.

Quality risk

False, repeated, or inconsistent rejections created friction for creators and slowed publishing. Disagreement between automated decisions and reviewers had to become visible, explainable, and actionable.

Operating loop

The workflow connects sampled game context, QA outcomes, explicit ground truth, calibrated failure modes, final-review comparison, and creator-feedback quality. Human judgment is an accountable control point, not a late-stage exception.

02 / Mandate

Mandate

Design the operating model

Create a calibration workflow that re-QAs completed game samples and turns QA, final review, and creator feedback into one continuous, evidence-based decision loop.

Make quality measurable

Define explicit ground truth, QA and final-review correctness, failure modes, agreement, and feedback actionability so quality is based on observable signals rather than reviewer sentiment.

Protect creator outcomes

Establish structured feedback and operational guardrails for review that preserve safety while testing whether a creator can act on the decision without guessing or re-discovering the issue.

evidence_task.log

Calibration checklist

Decision-validation, evidence-quality, and standard-alignment rules that make reviewer judgments comparable.

03 / Build

Build

1.Built role-scoped calibration sessions

Implemented repeatable sessions that load a completed game sample beside QA and final-review context, then collect structured calibration signals without free-text scoring.

2.Established explicit ground truth

Separated the expected platform decision from the QA decision, allowing the system to compute QA correctness and then diagnose why an incorrect decision occurred.

3.Defined QA and final-review guardrails

Added constrained failure-mode taxonomies for QA and final review so disagreement, policy misapplication, insufficient review, and tooling errors can be analyzed separately.

4.Closed creator feedback loops

Scored whether feedback is actionable and checks for a reproducible creator-facing formula, then saves immutable calibration records that can inform documentation and workflow updates.

evidence_action.log

Side-by-side calibration session

A sampled game runs next to QA and final-review context, keeping the source evidence available while calibration decisions are recorded.

Explicit ground truth

The calibrator selects the expected decision before evaluating whether QA reached it, separating decision correctness from reviewer sentiment.

Structured QA failure modes

When QA is incorrect, the session records a finite diagnosis such as false rejection, missed blocking issue, policy misapplication, or ambiguous standard.

Creator feedback quality

The tool scores feedback as fully, partially, or not actionable, then checks for a reproducible creator-facing formula.

04 / Outcomes

Outcomes

A measurable quality-control loop

Made QA incorrect decisions, false rejections, missed blocking issues, approval precision and recall, final-review correctness, QA-final agreement, and feedback actionability visible with counts and legacy-data exclusions.

More consistent decisions

Ground truth, structured failure modes, and separate QA and final-review comparisons create a shared basis for calibration, policy discussion, and reviewer guidance.

Creator feedback becomes a quality signal

Actionability checks turn creator feedback into a concrete review-quality measure, so improvements can target decisions that are safe, clear, and usable.

evidence_result.log

Final-review failure modes

Final-review correctness is computed from ground truth and outcome, then bounded by a separate diagnostic taxonomy.

Continue the evidence trail

From proof to role fit

Compare Astrocade AI QA Calibration Tool with adjacent systems, or carry its reviewed capabilities into an Adaptive Focus brief.

Human-in-the-loop AIEvaluation and calibrationModeration and QAOperational UX