October 2, 2026 · 5 min read
How to Summarize YouTube Videos with AI (Claude or OpenAI API)
Language models can't watch YouTube, but they read transcripts very well. A video summarizer is two calls: one to get the transcript, one to summarize it. This guide builds one in Python with the TranscriptYT API and an LLM API.
1. Get the transcript as LLM-ready text
Request format=md with paragraphs=true. Markdown output includes the title and channel as a header, and paragraph grouping merges short caption cues into readable blocks, which uses fewer tokens and gives the model better context than one line per cue.
Keep timestamps=true if you want the summary to cite moments in the video.
import os
import requests
def transcript_markdown(video: str) -> str:
res = requests.get(
"https://transcript-yt.com/v1/transcript",
params={"url": video, "format": "md", "paragraphs": "true", "timestamps": "true"},
headers={"Authorization": f"Bearer {os.environ['TRANSCRIPTYT_KEY']}"},
timeout=120,
)
res.raise_for_status()
return res.text2. Summarize it with an LLM
Pass the transcript with a specific instruction. Asking for a fixed structure (key points, takeaways, timestamps) gives far more useful results than "summarize this".
import anthropic
client = anthropic.Anthropic()
def summarize(video: str) -> str:
message = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{
"role": "user",
"content": "Summarize this YouTube transcript in 5 bullet points, "
"each with the [HH:MM:SS] timestamp where it is discussed.\n\n"
+ transcript_markdown(video),
}],
)
return message.content[0].text
print(summarize("https://youtu.be/dQw4w9WgXcQ"))Long videos
An hour of speech is roughly 10,000 to 15,000 words, which fits in the context window of current Claude and GPT models. For multi-hour streams or podcasts, split the paragraphs into chunks, summarize each, then summarize the summaries.
Videos without captions
Some videos have no captions at all. TranscriptYT transcribes those with AI speech-to-text automatically, so the same code works; the response's source field is speech_to_text and it costs 1 credit per started 3 minutes of audio.
Skip the code: use the MCP server
If you just want summaries in Claude Code, Claude Desktop, Cursor, or Codex, connect the TranscriptYT MCP server and paste a link into the chat. The agent fetches the transcript itself.