Industry Playbook · 2026

Media & Gaming

Event pipelines and content metrics at scale. Unify streaming and batch data without moving a single event row.

25 min read Runbooks Metadata only Apache 2.0

Executive summary

Connect stream + warehouse metadata into Metroflow. Certify DAU and watch-time in 60 days. Gate every schema change on blast-radius checks.

Media and gaming teams lose hours reconciling engagement numbers and hunting stale dashboards after Flink failures. Metroflow gives every team one searchable map of how live events flow into reports, with certified metrics and proof.

Faster impact analysis when a stream job fails
1Certified DAU across growth, content, and finance
<15mTo answer "which dashboards break if X fails?"
01

What broken lineage costs you

Typical patterns at streaming and gaming companies. Ranges, not guarantees.

Stream job fails at 3am

2–6 hours of Slack archaeology

With Metroflow: Blast-radius query in minutes. Stakeholders notified before standup.

DAU defined five ways

1–3 hours per exec review reconciling

With Metroflow: One certified metric ID. Rogue explores flagged.

Session model change

Surprise breakages in ML or board decks

With Metroflow: Pre-merge impact across streams, warehouse, and BI.

New analyst onboarding

2–3 weeks to first trusted lineage query

With Metroflow: Living graph day one. Company Brain in plain English.

Metadata only. Metroflow crawls schemas, job names, manifests, and dashboard definitions. Your video, gameplay, and event rows never leave your network.

02

Your stack, one graph

Live events and overnight batch jobs connected. Metroflow sits above the data path, not inside it.

When sessionize_stream fails, trace the blast radius: staging tables stop updating → dbt models break → Looker explores and board slides go stale. One query, full picture.

03

Where are you today?

Most teams land at L1 or L2. Target L4 in 90 days.

L1Siloed docsSeparate wikis for Kafka, dbt, and BI.
L2Partial lineageWarehouse lineage exists. Streams and dashboards disconnected.
L3Unified graphStreams, warehouse, and BI in one map.
L4Certified metricsDAU and watch-time owned and enforced.
L5Proactive opsPre-merge gates. Stale dashboards caught early.

Quick self-check

  • Trace DAU from dashboard tile to Kafka topic in under 5 minutes?
  • Growth and content share one definition of "active user"?
  • Flink failure → affected dashboards without opening Slack?
  • dbt PRs include cross-layer impact checks?
  • Named owner for watch-time and DAU?

0–2: Start Week 1 connect · 3–4: Certify metrics · 5: Add change gates

04

Choose your path

Media & gaming is not one stack. Pick the track closest to your business.

Mobile & PC gaming

Events: match_started, level_complete, purchase

Certify first: DAU, D1/D7 retention, session length, ARPDAU

Streaming & VOD

Events: episode_start, watch_heartbeat, subscribe

Certify first: Watch-time, concurrent viewers, completion rate

Live & esports

Events: stream_join, chat_message, concurrent_viewers

Certify first: Peak concurrency, avg watch, engagement rate

05

Who owns what

DAU fights are governance problems. Assign decision rights up front.

RoleOwnsOn Metroflow
Data platformKafka, Flink, AirflowRegister jobs. Blast-radius on failure. Own comms template.
Analytics engineerdbt models, session logicImpact before PR merge. Link models to topics.
BI leadLooker, exec dashboardsCertified metrics only. Deprecate rogue SQL.
Growth / contentMetric definitionsApprove DAU and watch-time. Self-serve via Brain.
Head of dataCertification sign-offGate changes. Weekly trust score on dashboards.
06

30 · 60 · 90 day rollout

A program with gates, not just a connector checklist.

Days 1–30

Connect & first win

  • Connect warehouse, dbt, orchestration, BI
  • Register Kafka topics + Flink jobs
  • First blast-radius query
Gate: One end-to-end lineage path documented
Days 31–60

Certify metrics

  • Publish DAU + watch-time specs
  • Migrate top 5 dashboards
  • Flag legacy explores
Gate: Exec dashboard on certified DAU only
Days 61–90

Operationalize

  • Pre-merge impact for session models
  • 3am runbook in on-call
  • Weekly stale-dashboard report
Gate: Zero surprise stale exec metrics
07

3am incident runbook

sessionize_stream is failing. Follow this timeline.

T+0 · Detect
Alert fires

Flink retrying or lag spiking. Open Metroflow, search job name.

T+5 min · Blast radius
Run impact query
"What dashboards break if sessionize_stream fails?"
T+15 min · Communicate
Notify stakeholders

Paste Slack template below. Tag BI lead and content ops.

T+30 min · Fix
Platform recovers

Restart, scale, or rollback Flink. Analytics monitors freshness.

T+60 min · Verify
Confirm dashboards fresh

Re-run lineage. Close with affected dashboard list.

Slack template

[INCIDENT] sessionize_stream failing (retry 3/5) Impact: staging.sessions_raw stale · fct_sessions not updating Affected: Content Performance (Looker), Exec Summary slide 4 Owner: @platform-oncall · ETA 30m BI: hold Monday content review until cleared Metroflow blast-radius: [paste link]
08

Metric certification pack

Copy into your governance doc. One definition. One owner. Full lineage.

DAU (Daily Active Users)

Certify first
Formula
Distinct user_id with ≥1 qualifying event per calendar day (UTC).
Event
session_start OR match_started OR episode_start
Exclusions
Bots, test accounts, sessions < 5 seconds.
Owner
Head of Growth + Analytics Eng
Lineage
raw_play_eventssessionize_streamfct_dau → Exec dashboard

Watch-time (minutes)

Certify second
Formula
Sum of watch_heartbeat intervals, 5 min gap cap.
Edge cases
Exclude background play after 60s idle. Pause > 3 min stops count.
Owner
Content analytics + Analytics Eng
Lineage
watch_heartbeatfct_watch_time → Content Performance

Session engagement rate

Certify third
Formula
% sessions with ≥2 meaningful events OR duration ≥3 min.
Why
Separates passive opens from real engagement. Stops DAU inflation.
Owner
Product analytics + Growth
09

Daily workflows

Four situations you will hit every week.

📊

Monday content review

  1. Check freshness

    Stale explores from weekend stream lag?

  2. Certified metrics

    Content Performance on watch-time ID?

  3. Ask Brain

    Pipeline delay or definition change?

🔀

Before a dbt PR

  1. Impact query

    Renamed columns in session models.

  2. Notify BI

    If certified explore affected.

  3. Attach diff

    Lineage diff in PR description.

⚖️

DAU dispute

  1. Open spec

    Formula, exclusions, owner.

  2. Trace lineage

    Dashboard tile → source topic.

  3. Flag rogues

    Migrate explores off legacy SQL.

🚀

Content launch

  1. Verify schema

    Event catalog before launch day.

  2. Pre-wire dashboard

    Certified engagement metrics.

  3. Baseline

    7-day watch-time for before/after.

10

Copy-paste queries

Company Brain or lineage search. Context included.

Stream job failure
"What dashboards break if the sessionize_stream job fails?"
Before exec review
"List BI explores that do not use the certified DAU metric."
End-to-end trace
"Trace DAU from raw_play_events to the executive summary dashboard."
Debug number movement
"Why did watch-time change last Tuesday: pipeline delay or definition change?"
Before schema change
"What depends on fct_sessions if we rename device_id?"
Content ops weekly
"Which Flink jobs feed tables used in the weekly content performance deck?"
11

Glossary

Plain English. "Why it matters" tells you when to care.

Kafka
High-speed bus for live events (clicks, plays, sessions).
Why: Kafka lag = every downstream report is late.
Flink
Turns raw events into session counts in real time.
Why: #1 cause of stale engagement dashboards.
Batch pipeline
Nightly jobs loading and transforming warehouse data.
Why: Batch and streams must agree on session logic.
Lineage
Map of where data came from and what it feeds.
Why: Without it, impact analysis is Slack archaeology.
Semantic graph
Metroflow's linked map of your entire stack.
Why: One search bar instead of five tools.
DAU
Unique users who did something meaningful in a day.
Why: The most argued metric. Certify once.
Watch-time
Total minutes users spent viewing content.
Why: Content and ad teams bet on this number.
Company Brain
AI that answers using your stack metadata, with citations.
Why: PMs query without a SQL ticket.
12

Outcomes checklist

Measure if the program is working.

Success metric90-day target
Time to answer "what breaks if X fails?"< 15 minutes
Exec dashboards on certified DAU100% of top 5
Cross-team DAU escalations / monthNear zero
Surprise stale metrics in reviewsZero

Ready to unify your stack?

Week 1: connect and run your first blast-radius query. Week 4: certify DAU. Week 8: operationalize the runbook.