Data Infra › Podcast Data

Data Infra · 02

Podcast Data

The only systematic monitor of investment podcasts and video shows: 850,000+ episodes across 605,846 feeds, distilled into company mentions where every row carries the verbatim quote, the resolved ticker, and a timestamped link into the publisher's own audio.

Every row passes the Bargo Verification Layer

Accurate · Reliable · Source‑backed

605,846
feeds tracked · 19,138 curated shows
444,220
company mentions, each with proof attached
~6 hrs
from episode publication to processed rows
93,319
private‑company mentions tracked too

What's inside

Who said what, where, and how often

Company mentions

Verbatim excerpt, resolved ticker and CIK, and a link to the source audio at the second it was said.

Show & venue history

Who hosts whom, how often, across years. The pattern no single episode can show.

Promotion flags

A weekly screen for small caps being promoted, joined with SEC filings and price. Rules frozen before outcomes.

Expert commentary

Vetted authority accounts and long-form interviews, transcribed and ticker-linked, image OCR and video included.

The data, not a description of it

A row you can check yourself

Every mention ships with the fields to verify it without trusting us: the publisher's own URL, the second, and the checks that already ran.

■ company_mention · illustrative format
ticker / cikNVDA · 0001045810 showsemiconductor investing podcast · episode 2026-08-12 t_start41:07 excerpt"HBM supply is the whole story for next year" audio_sourcepublisher's own feed · never a copy hosted by us evidence_gatepassed · quote located in transcript · security resolved substantivetrue · not a passing mention or list recital verifyfetch the feed, seek to 41:07, hear it yourself
Field names simplified for display. Full data dictionary ships with every evaluation snapshot.

Bargo Verification Layer · Accurate · Reliable · Source‑backed

Watch one mention survive the three gates

An episode drops. The AI hears a company being discussed. Here is what happens to that mention before it is allowed to reach you:

Gate 1 · The data

The quote exists, and so does the company

The excerpt must be located in the transcript, and the name must resolve to a real security. The AI can only choose from symbols that exist, or decline.

✓ quote located in transcript · NVDA resolved to a real listing
invented ticker for an ambiguous name · structurally impossible; ~28% of raw extractions die at this gate
Gate 2 · The method

Verifiable against the publisher, not against us

The row links to the publisher's own audio at the second. Nothing in the proof chain passes through Bargo.

✓ seek to 41:07 on the publisher's feed · hear it yourself
✗ episode deleted by the publisher · reported as unverifiable, never quietly dropped
Gate 3 · The answer

Numbers are computed by frozen rules

Attention series and promotion flags come from preregistered rules joined with filings and price. Flags are locked the moment they fire.

✓ flag fired by a rule frozen before outcomes · attention × filings × price
"this feels promoted" · an impression, never served

Only then does the mention ship, carrying its receipts: the excerpt, the second, the publisher's URL, the checks. That is the whole promise, across 850,000+ episodes and counting. See the full Verification Layer →

Why it compounds: publishers delete and move episodes constantly. Roughly one in eight episodes we hold is already gone from the open web. We have it; a competitor starting today cannot get it at any price, and that fraction grows every day we run.

Access

The same verified mentions, three ways

Ask it anything

Bargo Agent

Attention questions with receipts: who is talking, how often, and whether it is organic.

> Is the attention on this small cap organic or manufactured?
For AI agents & terminals

Bargo MCP / CLIs

Wire mentions into Claude, Cursor or your own agents, or pull straight into pandas.

$ bargo podcasts mentions NVDA --days 30 --format parquet
Programmatic

Bargo APIs

Mentions, venue history and promotion flags as clean endpoints, ticker and CIK on every row.

GET /v1/podcasts/mentions?ticker=NVDA GET /v1/podcasts/promotion-flags/latest

Who runs on it

Built for the names everyone is suddenly talking about

Hedge funds

Edge with receipts

What executives and promoters actually said, verifiable to the second, before it shows up anywhere else.

Short sellers & diligence

The promotion flag sheet

Small caps being promoted, flagged weekly with excerpts, timestamps and the EDGAR context: issuance, Form 144s, insider clusters.

Quants

Attention as a series

Coverage per ticker over years, snapshot-stable, joinable on ticker and CIK, with lookahead caveats documented, not hidden.

Coverage & delivery

The facts sheet

Coverage605,846 feeds tracked, 19,138 curated investment shows, 2,636 shows carrying mentions, 93,319 private-company mentions alongside listed names.
LatencyNew episodes processed within roughly 6 hours of publication.
History850,000+ episodes archived permanently, including recordings no longer available on the open web.
Measured qualityMedian 96% word overlap against independently re-transcribed audio; 97.1% of timestamps derived deterministically; full-table identifier checks. Known limitations documented per dataset, up front.
FormatsJSON over API, CSV / Parquet snapshots, digest-stamped.
IdentifiersTicker and CIK on every row, for day-one joins with your existing stack.
VerificationEvery row carries its source and its checks. How the Verification Layer works →
We would rather show you the data than describe it.
Request an evaluation snapshot