Video Ingest Lab · SpaceX audio benchmark
Audio Evidence Without a Video Pipeline
One official 60-minute SpaceX webcast. One audio-only fixture. Three speech routes. The result is not one universal winner—it is a cleaner routing contract.
Decision
Keep platform captions first. If captions are absent or inadequate, use ElevenLabs Scribe v2 when speaker separation and technical terminology matter most. AssemblyAI Universal-3.5 Pro produced the cleanest text relative to SpaceX’s edited transcript, but it merged eight analysts into one voice cluster.
Native video analysis is tabled for now. This SpaceX webcast displayed one static title card for the full hour; the useful evidence was entirely in the audio. Bob/Roxy remains unchanged. A benchmark is not a cutover.
The production route is audio-first
A video page is only the container. If the question is about speech, start with speech evidence already published by the source.
The route asks one question at each boundary: do we have permission, do source captions already answer the question, and—only if not—does a managed transcript justify uploading an audio-only copy? It does not send a static earnings webcast through a video model merely because the source file is MP4.
| Route | Use it for | Verdict |
|---|---|---|
| Platform captions | Public speech and source timing without another media upload | Primary |
| ElevenLabs Scribe v2 | Captionless permitted audio where speaker turns or exact domain terms matter | Default managed fallback |
| AssemblyAI Universal-3.5 Pro EU | Clean edited-style transcript, EU processing, or a prompt/keyterm-specific canary | Task-specific challenger |
| Native video analysis | Not part of the present test or production route | Tabled |
The SpaceX fixture
Q4’s registration page was not the whole archive. Public event metadata exposed the official recording and WebVTT captions without submitting registration data.[1]
The source was SpaceX’s Second Quarter 2026 webcast: 3,604.152 seconds of AAC audio at 44.1 kHz stereo. The original AAC stream was copied without re-encoding into one 72,706,500-byte M4A fixture, SHA-256 4db8a017…dbcdf. That exact file went to both managed providers.
The three routes were:
- Q4 WebVTT: the platform’s supplied captions and cue times; no additional upload.
- ElevenLabs Scribe v2: direct file upload, English, word timestamps, automatic diarization, audio events enabled, no speaker count and no keyterm prompt.[3]
- AssemblyAI Universal-3.5 Pro: EU upload, language detection, word timestamps, automatic diarization, formatting and punctuation, with Universal-2 only as the declared fallback; no speaker count and no keyterm prompt.[7]
Three routes, three different winners
AssemblyAI best matched the edited document. ElevenLabs best preserved domain language and speaker identity. Q4 still wins the route because it arrived with the source.
| Metric | Q4 captions | ElevenLabs Scribe v2 | AssemblyAI U-3.5 Pro |
|---|---|---|---|
| Normalized reference words | 7,796 | 7,796 | 7,796 |
| Normalized hypothesis words | 8,359 | 9,313 | 8,544 |
| Edited-relative WER | 13.40% | 23.79%* | 12.42% |
| Substitutions / deletions / insertions | 254 / 114 / 677 | 204 / 67 / 1,584* | 126 / 47 / 795 |
| Exact reference hit rate | 95.28% | 96.52% | 97.78% |
| Exact 19-term domain recall | 36.7% · 29/79 | 82.3% · 65/79 | 67.1% · 53/79 |
| Material figure groups | 11 pass · 1 partial | 12 / 12 pass | 12 / 12 pass |
| Automatic speaker clusters | None | 13 | 6 |
| Aligned speaker-word accuracy | Not available | 98.76% | 92.99% |
| Stable named turns | Not available | 38 / 39 · 97.44% | 31 / 39 · 79.49% |
| End-to-end completion | Already published | 55.89 s · 64.5× realtime | 51.59 s · 69.9× realtime |
| List-price estimate | No marginal upload | $0.2203 | $0.2303 including diarization |
| Deletion verification | Not applicable | Final GET 404 | Content-free tombstone |
* ElevenLabs’ 1,584 insertions dominate its WER. Its transcript retained substantially more spoken filler, repetitions, and false starts than the edited PDF. A lower WER against an edited document is not automatically a more faithful verbatim transcript.
What “accuracy” meant here
One aggregate WER would hide the important errors, so the audit used four layers:
- Speech-aware normalized WER: case and punctuation removed; written numbers, percentages, ordinals, currency,
GW,MHz, and10×-style forms normalized before alignment. - Reference hit rate: exact aligned reference words divided by reference words. This does not charge extra spoken fillers as missed reference content.
- Domain-term recall: exact occurrence recall over a fixed 19-term SpaceX glossary, with no provider keyterm prompt. It included names and terms such as Gwynne, Bret, Grok, EchoStar, Starmind, Rubin, NVL, Cursor, Claude, CapEx, Artemis, HLS, ARPU, and MVNO.
- Material semantic audit: twelve groups of financial, technical, unit, and chronology claims reviewed separately.
The decisive figure groups survived
Both managed routes preserved all twelve material groups. These included the opening $7.8B / 92% / $4.1B, $541M / $467M, and $3.5B / 191% / $1.2B figures; 78 launches and 1,041 tons; compute, capex, financing, cash and backlog; 20 GW versus roughly 15 GW; the $1T revenue timeline; 5 MHz and 65 MHz spectrum; the 10× then 100× capability claim; the roughly $600B mobile market; Artemis III, 2028, and Broadband V3 timing.
Q4 passed eleven groups and received one partial: in the power discussion it inserted “per week” after a repeated 20 GW phrase. The value remained present, but the added unit made the sentence materially ambiguous.
Terminology changed the ranking
ElevenLabs returned all 18 edited-reference occurrences of Grok, all six EchoStar occurrences, and the tested Starmind, Rubin, NVL, Cursor, Claude, CapEx, MHz, HLS, Artemis, Orion, ARPU, and MVNO terms. AssemblyAI rendered most Grok mentions as Groq; Q4 produced variants such as grog and GRC. All three commonly rendered Bret as Brett, and speaker-name spelling remained imperfect without keyterm prompting.
The speaker test changed the provider story
SpaceX’s edited transcript names 12 speakers; the webcast also contains an operator. ElevenLabs found 13 voice clusters. AssemblyAI found six.
For a reproducible speaker audit, the 39 named turns in the edited PDF were aligned to Q4 cue times through exact matching word blocks. Provider words inside those named-turn regions were then mapped one-to-one to the official speakers. Operator-only gaps absent from the edited PDF were excluded.
ElevenLabs
Thirteen total clusters—consistent with 12 named voices plus the operator. Within named regions it separated all 12 speakers, scored 98.76% speaker-word accuracy, and kept 38 of 39 turns on the globally mapped speaker label.
AssemblyAI
Six total clusters and only five inside named regions. It correctly separated the four management speakers, but collapsed eight analysts into one cluster. The 92.99% word-weighted score looks better than the actual attribution behavior because management delivered most of the words.
Q4 captions
No speaker labels. The captions remain useful speech evidence, but a downstream system must not invent attribution from unlabeled cues.
Identity boundary
Diarization produces anonymous clusters, not biometric identity. Names entered this audit only through alignment with the published edited transcript.
The previous clean single-presenter Tonbi canary caught AssemblyAI over-splitting one voice. This real earnings Q&A caught the opposite failure: under-clustering many short analyst turns. ElevenLabs remained the more reliable speaker-aware route across both shapes.
Both managed routes processed an hour in under a minute
AssemblyAI completed in 51.59 seconds end to end: 7.60 seconds to upload, 0.95 seconds to submit, and 40.79 seconds to process and poll. ElevenLabs’ direct request completed in 55.89 seconds. That is 69.9× and 64.5× realtime respectively.
At public pay-as-you-go rates observed August 6, Scribe v2 was $0.22 per hour.[4] AssemblyAI Universal-3.5 Pro was $0.21 per hour plus $0.02 per hour for speaker diarization.[7] The 60:04 fixture therefore estimates to $0.2203 and $0.2303. Q4 required no additional media upload or marginal STT purchase by this workflow.
Deletion was verified, not assumed
Both managed artifacts were deleted immediately after receipt. The providers expose different terminal states, so HTTP 200 alone is not the contract.
ElevenLabs
The first delete attempt returned 404 during eventual visibility. A bounded retry returned 200; the following GET returned 404. The artifact was accepted as non-retrievable only after that verification.[5]
AssemblyAI EU
DELETE returned 200. Verification GET also returned 200, but only as a content-free tombstone: text="Deleted by user.", a deleted marker in the audio URL, and no words or utterances.[9]
ElevenLabs documents Zero Retention Mode for eligible enterprise arrangements; otherwise explicit deletion remains part of the wrapper.[6] AssemblyAI’s retention documentation and EU routing remain relevant advantages, but the observed artifact still required semantic deletion verification.[8]
Bob/Roxy stays unchanged
This test improved the contract. It did not authorize deployment.
permitted video or audio
→ use platform captions when adequate
→ otherwise extract one audio-only fixture
→ choose managed STT by task:
ElevenLabs for speaker-aware evidence and terminology
AssemblyAI for a task-specific edited-text / EU canary
→ preserve timestamps and anonymous clusters
→ delete provider artifact; verify terminal state
→ grounded summary / Q&A from speech evidence only
Before activation, Bob still needs a Bob-owned credential, hard spending limits, bounded upload/poll/retry behavior, automated deletion alerts, and one genuinely difficult teleconference fixture with compression, interruptions, similar voices, and crosstalk. Explicit approval remains required for any route change.
What we are not building
- No native video-analysis branch for now. Earlier canaries remain historical evidence, not an active route.
- No local Whisper/WhisperX/pyannote service fleet for Bob/Roxy.
- No new scheduler, vector store, database, or GPU worker merely to summarize a webcast.
- No default voiceprint library; anonymous diarization is not identity proof.
- No DRM, CAPTCHA, registration, authenticated-session, or expiring-authorization bypass.
- No production cutover hidden inside a successful benchmark.
The heavier architecture remains documented in Video Ingestion Pipeline: the architecture we decided not to build. It is historical research, not the current plan.
Sources
- Q4 — SpaceX Second Quarter 2026 webcast archive.
- SpaceX — Second Quarter 2026 Edited Transcript.
- ElevenLabs — Scribe speech to text.
- ElevenLabs — API pricing.
- ElevenLabs — delete transcript API.
- ElevenLabs — Zero Retention Mode.
- AssemblyAI — Universal-3.5 Pro and diarization pricing.
- AssemblyAI — data retention and deletion.
- AssemblyAI — delete transcript API.