Video Ingest Lab · SpaceX audio benchmark

Audio Evidence Without a Video Pipeline

One official 60-minute SpaceX webcast. One audio-only fixture. Three speech routes. The result is not one universal winner—it is a cleaner routing contract.

Decision

Keep platform captions first. If captions are absent or inadequate, use ElevenLabs Scribe v2 when speaker separation and technical terminology matter most. AssemblyAI Universal-3.5 Pro produced the cleanest text relative to SpaceX’s edited transcript, but it merged eight analysts into one voice cluster.

Native video analysis is tabled for now. This SpaceX webcast displayed one static title card for the full hour; the useful evidence was entirely in the audio. Bob/Roxy remains unchanged. A benchmark is not a cutover.

Q4 captions: primary route ElevenLabs: speaker-aware fallback AssemblyAI: edited-text winner Native video: tabled Bob/Roxy: gated
60:04same source audio
12.42%best edited-relative WER
98.76%best speaker-word score
0production routes changed

The production route is audio-first

A video page is only the container. If the question is about speech, start with speech evidence already published by the source.

The route asks one question at each boundary: do we have permission, do source captions already answer the question, and—only if not—does a managed transcript justify uploading an audio-only copy? It does not send a static earnings webcast through a video model merely because the source file is MP4.

Audio-first evidence routing diagram: permitted media passes a rights gate, then uses platform captions when available or an audio-only managed speech-to-text fallback with verified deletion. Native video analysis is tabled.
The current production contract. Click for the vector version. Native video analysis is no longer an active lane.
RouteUse it forVerdict
Platform captionsPublic speech and source timing without another media uploadPrimary
ElevenLabs Scribe v2Captionless permitted audio where speaker turns or exact domain terms matterDefault managed fallback
AssemblyAI Universal-3.5 Pro EUClean edited-style transcript, EU processing, or a prompt/keyterm-specific canaryTask-specific challenger
Native video analysisNot part of the present test or production routeTabled

The SpaceX fixture

Q4’s registration page was not the whole archive. Public event metadata exposed the official recording and WebVTT captions without submitting registration data.[1]

The source was SpaceX’s Second Quarter 2026 webcast: 3,604.152 seconds of AAC audio at 44.1 kHz stereo. The original AAC stream was copied without re-encoding into one 72,706,500-byte M4A fixture, SHA-256 4db8a017…dbcdf. That exact file went to both managed providers.

The three routes were:

  1. Q4 WebVTT: the platform’s supplied captions and cue times; no additional upload.
  2. ElevenLabs Scribe v2: direct file upload, English, word timestamps, automatic diarization, audio events enabled, no speaker count and no keyterm prompt.[3]
  3. AssemblyAI Universal-3.5 Pro: EU upload, language detection, word timestamps, automatic diarization, formatting and punctuation, with Universal-2 only as the declared fallback; no speaker count and no keyterm prompt.[7]
The reference is not verbatim ground truth. SpaceX labels its PDF an “Edited Transcript” and warns that it can contain material errors, omissions, or inaccuracies.[2] Fillers, repetitions, false starts, operator lines, and some spoken wording are absent. Every WER below is therefore edited-transcript-relative WER, not absolute transcription error.

Three routes, three different winners

AssemblyAI best matched the edited document. ElevenLabs best preserved domain language and speaker identity. Q4 still wins the route because it arrived with the source.

Three-column SpaceX audio benchmark. Q4 captions win routing, AssemblyAI wins edited-transcript-relative WER, and ElevenLabs wins domain terminology and speaker separation.
Measured August 6, 2026 on one common 60:04 audio fixture. No keyterm prompts were supplied. Click for the vector version.
MetricQ4 captionsElevenLabs Scribe v2AssemblyAI U-3.5 Pro
Normalized reference words7,7967,7967,796
Normalized hypothesis words8,3599,3138,544
Edited-relative WER13.40%23.79%*12.42%
Substitutions / deletions / insertions254 / 114 / 677204 / 67 / 1,584*126 / 47 / 795
Exact reference hit rate95.28%96.52%97.78%
Exact 19-term domain recall36.7% · 29/7982.3% · 65/7967.1% · 53/79
Material figure groups11 pass · 1 partial12 / 12 pass12 / 12 pass
Automatic speaker clustersNone136
Aligned speaker-word accuracyNot available98.76%92.99%
Stable named turnsNot available38 / 39 · 97.44%31 / 39 · 79.49%
End-to-end completionAlready published55.89 s · 64.5× realtime51.59 s · 69.9× realtime
List-price estimateNo marginal upload$0.2203$0.2303 including diarization
Deletion verificationNot applicableFinal GET 404Content-free tombstone

* ElevenLabs’ 1,584 insertions dominate its WER. Its transcript retained substantially more spoken filler, repetitions, and false starts than the edited PDF. A lower WER against an edited document is not automatically a more faithful verbatim transcript.

What “accuracy” meant here

One aggregate WER would hide the important errors, so the audit used four layers:

  1. Speech-aware normalized WER: case and punctuation removed; written numbers, percentages, ordinals, currency, GW, MHz, and 10×-style forms normalized before alignment.
  2. Reference hit rate: exact aligned reference words divided by reference words. This does not charge extra spoken fillers as missed reference content.
  3. Domain-term recall: exact occurrence recall over a fixed 19-term SpaceX glossary, with no provider keyterm prompt. It included names and terms such as Gwynne, Bret, Grok, EchoStar, Starmind, Rubin, NVL, Cursor, Claude, CapEx, Artemis, HLS, ARPU, and MVNO.
  4. Material semantic audit: twelve groups of financial, technical, unit, and chronology claims reviewed separately.

The decisive figure groups survived

Both managed routes preserved all twelve material groups. These included the opening $7.8B / 92% / $4.1B, $541M / $467M, and $3.5B / 191% / $1.2B figures; 78 launches and 1,041 tons; compute, capex, financing, cash and backlog; 20 GW versus roughly 15 GW; the $1T revenue timeline; 5 MHz and 65 MHz spectrum; the 10× then 100× capability claim; the roughly $600B mobile market; Artemis III, 2028, and Broadband V3 timing.

Q4 passed eleven groups and received one partial: in the power discussion it inserted “per week” after a repeated 20 GW phrase. The value remained present, but the added unit made the sentence materially ambiguous.

The reference can be wrong too. In the lunar discussion, both managed systems returned “uncrewed” where the edited PDF and Q4 captions say “crude.” That disagreement was labeled a reference dispute rather than automatically charged against either managed provider.

Terminology changed the ranking

ElevenLabs returned all 18 edited-reference occurrences of Grok, all six EchoStar occurrences, and the tested Starmind, Rubin, NVL, Cursor, Claude, CapEx, MHz, HLS, Artemis, Orion, ARPU, and MVNO terms. AssemblyAI rendered most Grok mentions as Groq; Q4 produced variants such as grog and GRC. All three commonly rendered Bret as Brett, and speaker-name spelling remained imperfect without keyterm prompting.

The speaker test changed the provider story

SpaceX’s edited transcript names 12 speakers; the webcast also contains an operator. ElevenLabs found 13 voice clusters. AssemblyAI found six.

For a reproducible speaker audit, the 39 named turns in the edited PDF were aligned to Q4 cue times through exact matching word blocks. Provider words inside those named-turn regions were then mapped one-to-one to the official speakers. Operator-only gaps absent from the edited PDF were excluded.

ElevenLabs

Thirteen total clusters—consistent with 12 named voices plus the operator. Within named regions it separated all 12 speakers, scored 98.76% speaker-word accuracy, and kept 38 of 39 turns on the globally mapped speaker label.

AssemblyAI

Six total clusters and only five inside named regions. It correctly separated the four management speakers, but collapsed eight analysts into one cluster. The 92.99% word-weighted score looks better than the actual attribution behavior because management delivered most of the words.

Q4 captions

No speaker labels. The captions remain useful speech evidence, but a downstream system must not invent attribution from unlabeled cues.

Identity boundary

Diarization produces anonymous clusters, not biometric identity. Names entered this audit only through alignment with the published edited transcript.

The previous clean single-presenter Tonbi canary caught AssemblyAI over-splitting one voice. This real earnings Q&A caught the opposite failure: under-clustering many short analyst turns. ElevenLabs remained the more reliable speaker-aware route across both shapes.

Both managed routes processed an hour in under a minute

AssemblyAI completed in 51.59 seconds end to end: 7.60 seconds to upload, 0.95 seconds to submit, and 40.79 seconds to process and poll. ElevenLabs’ direct request completed in 55.89 seconds. That is 69.9× and 64.5× realtime respectively.

At public pay-as-you-go rates observed August 6, Scribe v2 was $0.22 per hour.[4] AssemblyAI Universal-3.5 Pro was $0.21 per hour plus $0.02 per hour for speaker diarization.[7] The 60:04 fixture therefore estimates to $0.2203 and $0.2303. Q4 required no additional media upload or marginal STT purchase by this workflow.

Deletion was verified, not assumed

Both managed artifacts were deleted immediately after receipt. The providers expose different terminal states, so HTTP 200 alone is not the contract.

ElevenLabs

The first delete attempt returned 404 during eventual visibility. A bounded retry returned 200; the following GET returned 404. The artifact was accepted as non-retrievable only after that verification.[5]

AssemblyAI EU

DELETE returned 200. Verification GET also returned 200, but only as a content-free tombstone: text="Deleted by user.", a deleted marker in the audio URL, and no words or utterances.[9]

ElevenLabs documents Zero Retention Mode for eligible enterprise arrangements; otherwise explicit deletion remains part of the wrapper.[6] AssemblyAI’s retention documentation and EU routing remain relevant advantages, but the observed artifact still required semantic deletion verification.[8]

Bob/Roxy stays unchanged

This test improved the contract. It did not authorize deployment.

permitted video or audio
  → use platform captions when adequate
  → otherwise extract one audio-only fixture
  → choose managed STT by task:
       ElevenLabs for speaker-aware evidence and terminology
       AssemblyAI for a task-specific edited-text / EU canary
  → preserve timestamps and anonymous clusters
  → delete provider artifact; verify terminal state
  → grounded summary / Q&A from speech evidence only

Before activation, Bob still needs a Bob-owned credential, hard spending limits, bounded upload/poll/retry behavior, automated deletion alerts, and one genuinely difficult teleconference fixture with compression, interruptions, similar voices, and crosstalk. Explicit approval remains required for any route change.

What we are not building

  • No native video-analysis branch for now. Earlier canaries remain historical evidence, not an active route.
  • No local Whisper/WhisperX/pyannote service fleet for Bob/Roxy.
  • No new scheduler, vector store, database, or GPU worker merely to summarize a webcast.
  • No default voiceprint library; anonymous diarization is not identity proof.
  • No DRM, CAPTCHA, registration, authenticated-session, or expiring-authorization bypass.
  • No production cutover hidden inside a successful benchmark.

The heavier architecture remains documented in Video Ingestion Pipeline: the architecture we decided not to build. It is historical research, not the current plan.

Sources

  1. Q4 — SpaceX Second Quarter 2026 webcast archive.
  2. SpaceX — Second Quarter 2026 Edited Transcript.
  3. ElevenLabs — Scribe speech to text.
  4. ElevenLabs — API pricing.
  5. ElevenLabs — delete transcript API.
  6. ElevenLabs — Zero Retention Mode.
  7. AssemblyAI — Universal-3.5 Pro and diarization pricing.
  8. AssemblyAI — data retention and deletion.
  9. AssemblyAI — delete transcript API.