Behind the curtain
How it works
The panel hears you
While you pitch, your microphone audio is cut into 3 second segments and each one is transcribed live through Speko's speech-to-text API (/v1/transcribe). Segments are processed and discarded. Nothing is recorded, nothing is stored.
The interruptions
The triggers are plain heuristics running on your live transcript: filler bursts, thirty seconds without a number, monologuing, buzzwords, and silence. The panel only ever reacts to what you actually said. It knows nothing about you or your company, and the roast targets delivery only.
Most interruption lines were synthesized ahead of time with Speko's /v1/synthesize voices. A couple per run are synthesized live from what you said, and the HUD shows the measured server time for those in milliseconds. That number is real, per request, not a marketing claim.
The score
Survival time, interruptions, filler count, pace, and buzzwords feed a deterministic score. The percentile compares you against a seed distribution that recalibrates as more runs come in.
The stack
Speech-to-text, the three voices, and the live synthesis all run on the Speko API. The barge-in speed you just experienced is the product.