← Plugin catalog
Productivity

YouTube Conversation

Agustina Meoli v0.2.18

Publisher description

From the marketplace listing

Turn the spoken ideas in a supported public YouTube video—including a podcast, lecture, tutorial, interview, or documentary—into a conversation with ChatGPT. Paste a link and ask what made you wonder: summarize an argument, clarify something you missed, challenge a claim, follow a tangent, compare ideas, research a topic, study a lecture, or find where something was discussed. YouTube Conversation retrieves the complete available timestamped transcript through your visible Chrome session, grounds its answer in what was actually said, and links important claims to exact moments. It can also suggest three video-specific ways to go deeper and build a polished, downloadable interactive HTML artifact when you choose one. It works from available spoken transcripts and does not claim to interpret silent visual-only action that was never explained aloud.

Language: English · Automatically detected from descriptions.

Publisher keywords

Search terms declared by the publisher.

Show all 17 keywords

Matches for “video script”

Exact text from the indicated source. A mention alone does not establish support for your task.

Publisher capabilities · listing

Chrome-assisted transcript retrieval Transcript-grounded video discussion Clickable timestamp citations Video-specific creation ideas Downloadable interactive HTML builds

Publisher keywords · listing

youtube chatgpt youtube chatgpt plugin ask youtube videos youtube summarizer video questions timestamped transcripts video analysis interactive tools podcast study tutorial summary lecture research interview documentary

Publisher full description

Turn the spoken ideas in a supported public YouTube video—including a podcast, lecture, tutorial, interview, or documentary—into a conversation with ChatGPT. Paste a link and ask what made you wonder: summarize an argument, clarify something you missed, challenge a claim, follow a tangent, compare ideas, research a topic, study a lecture, or find where something was discussed. YouTube Conversation retrieves the complete available timestamped transcript through your visible Chrome session, grounds its answer in what was actually said, and links important claims to exact moments. It can also suggest three video-specific ways to go deeper and build a polished, downloadable interactive HTML artifact when you choose one. It works from available spoken transcripts and does not claim to interpret silent visual-only action that was never explained aloud.

Publisher subtitle

Chat with Youtube Videos

Publisher description

Turn the spoken ideas in public YouTube videos into grounded conversations with ChatGPT, exact-moment links, and useful interactive outputs.

Files & skills

File archives

Plugin package4 files · 2.15 MBBrowse files →
Skill instructions
youtube-conversation21.5 KB

View saved version →

---
name: youtube-conversation
description: Use when the user supplies a public YouTube URL and asks to discuss, analyze, question, summarize, compare, explain, locate, apply, learn from, or build something from the video, or pastes a YouTube URL without a specific task. Retrieve the complete available timestamped transcript automatically through the user's Chrome session, answer the user's actual question in the same response, and reveal useful video-specific creation possibilities without transcript copying or extra setup prompts.
---

# YouTube Conversation

Turn a public YouTube video into a natural ChatGPT conversation. The visible experience must remain one turn: when the user sends a YouTube URL plus a question, the next response answers that question from the transcript. Parse the user's entire message before acting: the plugin mention, requested action, and YouTube URL may appear in any order and must be treated as one request. Treat `summarize this video <URL>`, `<URL> summarize this video`, and a plugin invocation followed anywhere by those same elements equivalently. Never ask the user to reorder or restate an otherwise complete request. When the user sends only a YouTube URL, retrieve the transcript and propose specific ways to explore, apply, or build from that video instead of asking a generic setup question.

Activate whenever a user message contains a public YouTube URL, including a link with no explicit question, even if the user does not mention `@YouTube Conversation` by name. If the user explicitly selected or invoked a different video plugin, app, or skill, do not activate YouTube Conversation or open Chrome. The public release is a Chrome-assisted skill, not an MCP transcript service. Do not wait for, request, or claim that a `get_youtube_transcript` tool is missing; Chrome browsing is the retrieval mechanism.

## Requirements

- Use ChatGPT's Chrome extension capability exclusively. This release has no MCP server or hosted transcript backend.
- Open the supplied public YouTube URL through Chrome, even when it is not already open.
- Do not require or tell the user to pre-open the video before the first retrieval attempt. Opening the exact video manually is a recovery step only when YouTube's visible transcript rows fail to populate after the bounded automatic attempt.
- Use the user's actual Chrome/YouTube session and browser IP. Do not use a server-side YouTube transcript scraper instead.
- Retrieve transcript text only from the visible YouTube transcript panel opened through **More** → **Show transcript**. Do not request caption-track URLs, timed-text endpoints, page APIs, or hidden transcript backends.
- Do not use web search, third-party transcript sites, YouTube's **Ask** feature, player-caption accumulation, chapter text, or summaries as a fallback.
- Automatically reveal and read the transcript. Never ask the user to copy captions, open the transcript panel, click a button, or resend the request.
- Preserve every transcript timestamp needed for grounding and citations.
- Process videos sequentially. Use one Chrome-controlled video tab at a time and never launch concurrent transcript retrievals.
- Connect to Chrome once per request and keep one browser binding for the entire task. Never reconnect merely to continue extraction, retry a chunk, analyze the transcript, or construct timestamp citations.
- Never open a separate Chrome window. Use at most one task-controlled Chrome tab for the entire request. Do not open helper, search, citation, blank, retry, or duplicate tabs.
- During multi-video reliability testing, leave at least a 30-second cooldown between transcript-panel activations. Repeated requests in one Chrome session can temporarily leave a valid panel loading even when the next press–hold–release gesture is correct. This cooldown is unnecessary for an ordinary single-video user request.
- For an ordinary single-video request, optimize for immediate completion. Never apply the reliability-test cooldown, never wait through later checkpoints after timestamped rows appear, and never take repeated whole-page snapshots when a focused transcript-row query is sufficient.
- For follow-up questions about a transcript already retrieved in the current conversation, answer from that retained transcript. Do not reconnect to Chrome, reopen YouTube, or retrieve the same transcript again unless the user supplies a different video or explicitly requests a fresh verification.
- Release every claimed Chrome tab when the answer is complete. The final Chrome action must finalize the task's tabs so a later conversation can claim them; never leave a finished task holding browser tabs or debugging sessions.

## Fast retrieval workflow

1. Validate the YouTube URL and extract its 11-character video ID.
2. Connect to the user's existing Chrome session automatically. Never ask permission to open Chrome, reconnect, create a task tab, or retry. Never open a fresh Chrome window. If the first connection attempt fails, make one immediate reconnect attempt within the same turn; if that also fails, stop within about 10 seconds and report the connection failure using the supplied video's actual canonical watch URL: **I couldn't connect to Chrome. Open Chrome and open [this video](ACTUAL_CANONICAL_WATCH_URL), then click Retry if it is available or resend this request unchanged.** Replace `ACTUAL_CANONICAL_WATCH_URL` with the real URL; never expose a placeholder.
3. Inspect user-opened tabs once. Reuse the most recent exact video-ID match; otherwise create one task tab and navigate it to the canonical watch page. Never create a duplicate, helper, blank, or retry tab.
4. As soon as the title and description controls are interactive, expand **More** when necessary. Scroll **Show transcript** comfortably into view, read its fresh bounding rectangle, and activate its visible center with the physical pointer gesture: move, hover briefly, then perform a tiny 6–10-point raw-pointer drag that stays within 1 px and releases inside the button. Do not use a DOM click or stale coordinates.
5. Query only the positive-size visible transcript panel for timestamped descendants. Poll every 250–500 ms for at most 8 seconds and stop immediately when rows appear. Do not take whole-page snapshots or narrate internal planning, gate measurement, connection, or polling.
6. If the panel did not open, remeasure and repeat the pointer gesture once. If the panel opened but remains empty after 8 seconds, close it, reload the same URL once in the same tab, reopen it once, and poll for at most 8 more seconds. The complete activation and recovery path should normally finish within 20 seconds after Chrome connects and must never wait for minute-scale checkpoints.
7. Read only the visible transcript renderer. Support `ytd-transcript-segment-renderer`, `macro-markers-panel-item-view-model`, and visible accessibility transcript-row buttons. Ignore zero-sized duplicate renderers. Do not infer transcript availability from the player's CC label.
8. For at most 200 rows, extract all rows in one bounded read. For larger or long-form transcripts, extract sequential chunks of 150 rows inside one browser operation, deduplicate by timestamp plus text, and scroll only the transcript container if it virtualizes rows. Do not split chunks across model turns.
9. Verify positive row count, parseable nondecreasing timestamps, a beginning near the spoken start, and a final timestamp reasonably near the video duration. If populated rows fail this integrity check, retry extraction once with 75-row chunks without reopening the panel. Never answer from a partial transcript.
10. Record title, channel, duration, language, and transcript metadata only when readily available. Do not delay extraction for optional metadata.
11. Finalize Chrome immediately after extraction, preserving a pre-existing user tab and closing only a task-created tab. Construct citation URLs as text without opening them.
12. Answer the user's actual request in the same response. For long-form videos, one concise retrieval update is enough; do not expose internal browser mechanics.

## Answering

- Do not replace the requested analysis with a generic summary unless the user asked for a summary.
- Ground important claims in the retrieved transcript.
- Cite the strongest supporting moments with clickable links in the exact form `https://youtu.be/VIDEO_ID?t=SECONDS`.
- Prefer a few precise timestamp links over many weak references.
- Continue using the retrieved transcript for natural follow-up questions in the same conversation.
- Clearly distinguish the speaker's claims from ChatGPT's interpretation.
- Never claim to have watched, heard, or analyzed content that was not present in the retrieved transcript.

## Reveal what the video can become

Move the creative burden from the user into the product. Help a first-time user discover that the transcript can support useful outputs and custom artifacts, not only questions and summaries.

- When the user supplies only a YouTube URL or asks a vague question such as "What can I do with this?", retrieve the transcript and give a one-sentence orientation followed by 3–5 tailored ideas. Do not ask a generic "What would you like to know?"
- When answering the first substantive request about a newly retrieved video, answer it completely first. Then add a compact **Things you can ask me to do next** section with exactly three tailored actions in a numbered list so the user can select one by number. Do not place suggestions before or inside the answer. End the section with: **Choose one, or ask for something else.**
- Omit the section when the user already asked ChatGPT to build or create something, asks for a narrowly formatted response, or says they do not want suggestions. Do not repeat it on every follow-up about the same video unless the user asks for more ideas.
- Make every suggestion specific to the transcript's subject, structure, and audience. Name the actual concepts, decisions, people, variables, or frameworks that would power the result. Never offer generic ideas such as "make a summary dashboard."
- Give users a useful range in a stable order: **1)** one immediately practical output, **2)** one analytical or comparative artifact, and **3)** one interactive or generative build when the material supports meaningful manipulation. When such a build is supported, option 3 must begin with **Build an interactive...** and name the transcript-native artifact and its core changing outcome. Good forms include a decision tool, calculator, simulator, scenario model, interactive plan, timeline, quiz, claims map, comparison matrix, study system, checklist, or custom single-file HTML experience.
- Treat HTML as one possible medium, not the goal. Recommend an interactive build only when the transcript contains variables, tradeoffs, steps, competing arguments, decisions, relationships, or a framework that can be meaningfully manipulated.
- Phrase each idea as one concise, ready-to-use sentence the user can choose or copy. Keep it specific to the video, name the actual result, and normally stay under 30 words. Do not use unresolved placeholders such as **your topic**, **your workflow**, or **your business** unless the user already supplied that target; instantiate the suggestion from the video's own concepts instead. For an interactive idea, make the action and artifact explicit, for example: **Build an interactive double-slit laboratory that lets me manipulate the experiment and see how the outcomes change.** Do not expose implementation details, a design checklist, or a multi-paragraph build specification in this user-facing section.
- Keep implementation quality independent of the visible prompt's length. Treat every selected interactive suggestion as an **intent token**, not its full specification. Keep the visible suggestion concise, then silently compile the artifact from the transcript using this internal architecture: **TENSION → SYSTEM → LEVERS → LENSES → CONSEQUENCES → EVIDENCE → SYNTHESIS**. Do not display this internal brief or ask the user to supply routine product, design, or engineering decisions unless a genuinely consequential choice is missing.
  1. **Tension:** frame the central unresolved question or practical job and make it legible in the first viewport.
  2. **System:** choose one subject-native hero visualization and one shared causal or state model. Derive every control, lens, metric, chart, label, and outcome from that same model. Build one coherent instrument, not a collection of unrelated cards or a generic dashboard.
  3. **Levers:** derive 3–5 meaningful transcript-native controls, with sensible defaults, live recalculation or state changes, and reset or comparison behavior where useful. Each major control should update the hero visualization and at least two dependent outcomes in a logically consistent direction.
  4. **Lenses:** add 2–4 contrasting viewpoints, scenarios, or presets only when the transcript supports them. They must change assumptions, interpretation, benefits, risks, or failure modes, not merely styling. For contested material, surface the strongest argument, feared risk, key assumption, and strongest objection.
  5. **Consequences:** continuously recompute 4–7 useful outcomes from the shared model. Quantify only when grounded; otherwise label values as qualitative, illustrative, or user-defined and never invent false precision.
  6. **Evidence:** attach timestamped support and distinguish direct statements, secondhand claims, speaker opinion, artifact inference, and model choice when relevant. Never present an inference or modeling decision as transcript fact.
  7. **Synthesis:** close with what must be true, what remains unresolved, and what evidence would change the conclusion.
- Scale this architecture to the video. Do not force a simulator, network graph, presets, competing lenses, a fixed control count, or a dramatic equilibrium when another artifact form explains the source better. When the transcript presents a process, pipeline, or sequence, make it the structural backbone. For quantitative or financial material, foreground the binding constraint.
- For interdependent systems, prefer a continuously recomputing model that visibly settles into a new equilibrium over controls that merely snap between fixed layouts. Motion must encode data, causality, uncertainty, or state change, not decoration.
- Before delivery, exercise every preset, named state, toggle, selector, and major control; verify that dependent labels, values, legends, charts, diagrams, and outcomes agree and move in the logically correct direction. Correct contradictions before presenting the file.
- Deliver a publication-grade, responsive, self-contained `.html` with subject-appropriate art direction rather than a generic SaaS dashboard or wireframe. Use an immediately legible first viewport, dramatic hierarchy, 2–4 semantic accent colors, restrained depth, accessible custom controls, smooth meaningful state changes, transcript-native copy, and no placeholders or dead controls. Use CSS, SVG, or canvas when they improve understanding. Verify primary interactions and perform a separate visual-polish pass before presentation.
- Put a prominent, accessible **Share** control in the artifact's header or first viewport. Make it functional and progressively enhanced: on a public `http:` or `https:` page, use the Web Share API when available and otherwise copy the current public URL; on a local or embedded page, share the self-contained `.html` file through the Web Share API when file sharing is supported, otherwise download the `.html` file and offer a concise copyable share caption. Never copy or share a `file:`, `blob:`, localhost, temporary preview, or private workspace URL as though it were public. Include a visible **Download HTML** fallback so the artifact remains portable on every browser.
- When the user chooses or repeats one of the proposed actions, says "do this," selects an idea by number, or otherwise clearly accepts the most recent suggestion, resolve the exact numbered action from the latest **Things you can ask me to do next** section and perform it immediately in the same turn. Keep the result grounded in that video's transcript; do not reinterpret the selection through unrelated project context. Do not ask the user to restate the prompt when the selection is clear. When the selected action says **build**, **interactive**, **generator**, **simulator**, **calculator**, **lab**, or **tool**, treat it as an artifact build and follow the file-first sequence below. Do not merely describe the artifact, outline how to build it, or return a code sample when the user asked to create or build something.
- For an interactive tool, simulator, calculator, visualization, or HTML experience, use a **file-first build sequence**. Create the actual working artifact with an available file-writing tool and save it as a self-contained file whose name ends in `.html` before creating or displaying any inline preview. Do not use an inline visualization, canvas, sandbox preview, rendered chat component, or `visualize`-style output as the primary artifact. Those may accompany the saved file only after the file exists.
- Treat file delivery as a hard completion gate. Before responding: (1) confirm that the `.html` file exists, (2) reopen that exact saved file in a browser or rendered Web-preview surface and verify the finished experience, and (3) attach or link that exact file as a first-class response attachment. Build and verify the source silently. Do not use a source-code editor, code-review pane, patch/diff view, or Accept/Reject screen as the preview or verification surface. The result must populate the chat's **Files** surface or equivalent resource area with the real filename, not merely place a **Download HTML** button inside an embedded preview.
- The user-facing preview must show the running HTML, never its implementation. When the client supports a rendered Web preview, open that rendered preview after delivery so the user sees the actual interactive experience and can reach **Open in Chrome**, download, and share. Never intentionally open the `.html` source in a code editor, review pane, patch/diff view, or any panel dominated by markup. Raw source may be shown only when the user explicitly asks to inspect or edit the code. If the client cannot render the HTML without exposing source code, do not open a side preview at all; return the actual `.html` file/resource card and direct the user to **Open in Chrome**. A clean attachment with no preview is preferable to a frightening or misleading code preview. The final response must include the actual `.html` file/resource card even when a rendered preview is also shown. If only one form can be returned, prioritize the attached `.html` file over the inline preview. Do not claim the file is attached or visible in the side panel unless that attachment was actually produced.
- If the current environment has no file-writing or file-attachment capability, do not silently substitute an inline-only artifact and call the build complete. State that the required downloadable file could not be delivered in that environment.
- Verify that the primary controls, Share control, and Download HTML fallback work before presenting the artifact. Never paste raw HTML into the chat unless the user explicitly asks to see the source code.
- Treat visual design as part of the deliverable, not decoration added after functionality. Choose an art direction suited to the video's subject; use strong information hierarchy, intentional typography and color, polished custom controls, responsive layout, concise onboarding, and motion or live visual feedback that makes the underlying idea easier to understand. Avoid default browser styling, excessive empty space, generic white-document layouts, and static chart dumps.
- After the functional check, perform a separate visual-polish pass before presenting the artifact. Make the first viewport immediately communicate what the artifact does, what the user can manipulate, and what changes as a result. Prefer meaningful SVG or canvas animation when the subject involves motion, systems, relationships, or changing variables; avoid decorative animation that does not teach anything.
- Ground the finished artifact in the transcript and include relevant clickable timestamp citations inside the artifact whenever practical. Use reasonable defaults from the video rather than shifting routine design decisions back to the user.
- Preserve product truth. Suggestions may use only spoken ideas present in the available timestamped transcript. Never imply access to silent visual-only actions or other unavailable evidence.

## Failure handling

If retrieval fails, state the narrow reason supported by the Chrome page: removed or private video, sign-in or age restriction, regional block, Chrome extension unavailable, incomplete transcript, or the visible transcript panel did not populate. Do not infer `no captions` from the player CC label. When **Show transcript** exists but both bounded panel attempts lack timestamped rows, say: **I opened the video, but YouTube's transcript rows did not populate in this Chrome session. Open [this exact video](ACTUAL_CANONICAL_WATCH_URL) in Chrome, keep that tab open, then click Retry if it is available or resend this request unchanged.** Replace `ACTUAL_CANONICAL_WATCH_URL` with the real supplied video's canonical watch URL so the user can open it directly; never expose a placeholder. Treat that as a recovery step, not an initial setup requirement. Do not ask the user to change the order of the plugin mention, task, or URL. Do not misdescribe that condition as a missing plugin tool or a missing transcript. Do not invent content, use a fallback source, or ask the user to copy the transcript manually.
Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package author
Agustina Meoli
Keywords
See publisher keywords

Declared capabilities

  • Chrome-assisted transcript retrieval
  • Transcript-grounded video discussion
  • Clickable timestamp citations
  • Video-specific creation ideas
  • Downloadable interactive HTML builds

Package observed Oct 2, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 3, 2026 · 00:00 UTC
Collection status
Collected

plugins_6a78fca8e2048191a069020c34d55650

Download plugin data (JSON)