All guides

October 2, 2026 · 6 min read

YouTube Transcripts for RAG: Load Videos into LangChain and Vector Search

Video is full of knowledge that never reaches your RAG pipeline because it isn't text. Transcripts fix that, and timestamped segments let every answer link to the exact moment it came from. This guide loads YouTube transcripts into LangChain documents ready for any vector store.

Fetch paragraph-sized chunks with timestamps

Raw caption cues are a few words each, too small to embed well. Request JSON with paragraphs=true and the API merges cues into paragraph-length segments on natural speech pauses. Each segment keeps its start time.

import os
import requests

def fetch_segments(video: str) -> dict:
    res = requests.get(
        "https://transcript-yt.com/v1/transcript",
        params={"url": video, "paragraphs": "true"},
        headers={"Authorization": f"Bearer {os.environ['TRANSCRIPTYT_KEY']}"},
        timeout=120,
    )
    res.raise_for_status()
    return res.json()

Turn segments into LangChain documents

Store the video ID, title, and start time as metadata. The start time becomes a deep link, so a retrieved chunk can cite youtube.com/watch?v=ID&t=SECONDS.

from langchain_core.documents import Document

def load_youtube(video: str) -> list[Document]:
    data = fetch_segments(video)
    return [
        Document(
            page_content=seg["text"],
            metadata={
                "video_id": data["videoId"],
                "title": data["title"],
                "start": seg["start"],
                "source": f"https://www.youtube.com/watch?v={data['videoId']}&t={int(seg['start'])}s",
            },
        )
        for seg in data["segments"]
    ]

docs = load_youtube("https://youtu.be/dQw4w9WgXcQ")
# vector_store.add_documents(docs)

Index a whole channel or playlist

Loop over video IDs and load each one. Results are cached, so re-indexing a video you've already fetched is fast. Failed requests are never billed, and you can skip videos that return NO_CAPTIONS or VIDEO_UNAVAILABLE.

Tips for better retrieval

A few choices make a large difference in answer quality:

  • Prepend the video title to each chunk before embedding so short chunks keep their context.
  • Use translate_to to index multilingual videos in one language, or store the language as metadata and filter on it.
  • Return the source link with every answer so users can jump to the moment in the video and verify it.

Try the TranscriptYT API

Get transcripts as JSON, text, SRT, or VTT. 100 free credits, no card required.

More guides