Pipeline running

Stop the pipeline?

Every stage pauses before its next step and nothing is published. It's never the wrong call when something looks off.

Reports

Is the system getting better, is it earning, and where is everything right now?

Not ready to scale up yet

2 of 4 deciding tests passedBlocking: Blind voice discrimination at chance level; Gate calibration — Cohen's kappa vs human labels

Tests that decide whether we scale up2 of 4 passed
  • Voice backtest — G2 threshold derived from real varianceMeasured live
    witty n=25, warm n=24
    Pass when: at least 20 pieces per voice
    Passed
  • Blind voice discrimination at chance levelMeasured live
    sara 72.5%, sara 51.2%
    Pass when: accuracy within 2 standard errors of 50%
    Failed
  • Gate calibration — Cohen's kappa vs human labelsMeasured live
    G2 k=0.402, G3 k=0.801, G4 k=0.858, G6 k=0.661, G7 k=0.867, G8 k=1.000
    Pass when: kappa >= 0.60 on every scored gate
    Failed
  • Cost within 30% of model, caching confirmed workingMeasured live
    $0.2116/piece
    Pass when: within 30% of $0.2550 modelled
    Passed
Other safety tests2 of 5 passed
  • Kill-gate audit — killed pieces were genuinely thinRecorded by hand
    18 of 50 reviewed; 15 agreed the kill was right
    Pass when: killed pieces were genuinely thin
    [demo] audit under way
    In progress
  • Audience-swap test at scaleRecorded by hand
    Not run yet
  • Dry-run publish — schema, links, idempotencyRecorded by hand
    200 pieces to staging; replayed publish produced no duplicate
    Pass when: schema valid, links resolve, idempotency holds
    [demo]
    Passed
  • Failure injection — resume, backoff, canary containmentRecorded by hand
    worker kill resumed mid-stage; canary held a bad profile to 10 pieces
    Pass when: all six injections behave
    [demo]
    Passed
  • Canary in production — 72h hold, no manual actionRecorded by hand
    Not run yet