Reports
Is the system getting better, is it earning, and where is everything right now?
Not ready to scale up yet
2 of 4 deciding tests passedBlocking: Blind voice discrimination at chance level; Gate calibration — Cohen's kappa vs human labels
Tests that decide whether we scale up2 of 4 passed
- PassedVoice backtest — G2 threshold derived from real varianceMeasured livewitty n=25, warm n=24Pass when: at least 20 pieces per voice
- FailedBlind voice discrimination at chance levelMeasured livesara 72.5%, sara 51.2%Pass when: accuracy within 2 standard errors of 50%
- FailedGate calibration — Cohen's kappa vs human labelsMeasured liveG2 k=0.402, G3 k=0.801, G4 k=0.858, G6 k=0.661, G7 k=0.867, G8 k=1.000Pass when: kappa >= 0.60 on every scored gate
- PassedCost within 30% of model, caching confirmed workingMeasured live$0.2116/piecePass when: within 30% of $0.2550 modelled
Other safety tests2 of 5 passed
- In progressKill-gate audit — killed pieces were genuinely thinRecorded by hand18 of 50 reviewed; 15 agreed the kill was rightPass when: killed pieces were genuinely thin[demo] audit under way
- Not run yetAudience-swap test at scaleRecorded by hand
- PassedDry-run publish — schema, links, idempotencyRecorded by hand200 pieces to staging; replayed publish produced no duplicatePass when: schema valid, links resolve, idempotency holds[demo]
- PassedFailure injection — resume, backoff, canary containmentRecorded by handworker kill resumed mid-stage; canary held a bad profile to 10 piecesPass when: all six injections behave[demo]
- Not run yetCanary in production — 72h hold, no manual actionRecorded by hand