Assurance

Quality

Quality here means what the validation pipeline has actually established, and nothing more. Checks that have not been built are named as such rather than quietly implied.

100%

Accepted for annotation

1 of 1 judged items passed every stage that ran.

0%

Flagged for human review

Usable but imperfect โ€” clipping, long silences. Warned, never auto-dropped.

0%

Rejected

Unusable audio: silent, truncated, or below the minimum sample rate.

Validation stages

Which validation stages have run against this corpus
Stage State What it establishes
Technical checks Running Duration, sample rate, clipping and silence, measured per asset.
Duplicate detection Not built Near-duplicate recordings and texts, so one submission cannot be paid for twice.
Language identification Not built Confirms the recording is in the language the task asked for. A mismatch warns; it never auto-rejects, because code-switching is expected and research-relevant.
Fraud signals Not built Replayed audio, implausible capture patterns, device-clock skew.

The model measures; this platform judges. Thresholds live in configuration here, not in the model, so swapping the checker changes a version string and nothing about what counts as acceptable. Currently reporting: technical/1.0.0. When the model is unreachable, jobs wait โ€” a model that is down never rejects a contributor's work.