VideoChaptersConverter
Automatically segment videos into topic-based chapters with titles, timestamps, and summaries by analyzing both the visual content and spoken dialogue. Produces chapter markers compatible with YouTube, Vimeo, and custom video players.
How It Works
Upload a video file or provide a URL to the Mixpeek API.
Audio is transcribed and visual scene changes are detected simultaneously.
Topic modeling on the transcript identifies semantic shift points.
Visual cues (title cards, slide transitions, scene changes) are correlated with topic boundaries.
An LLM generates chapter titles and summaries from the combined audio-visual context.
Code Examples
import os, requests
API = "https://api.mixpeek.com"
H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
"X-Namespace": os.environ["NAMESPACE_ID"]}
# 1. a bucket, with a schema that declares the field you will send
bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
"bucket_name": "video-inputs",
"bucket_schema": {"properties": {"video": {"type": "video"}}},
}).json()
# 2. land the file as an object. the URL goes in data, on the blob
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
"key_prefix": "run-1",
"blobs": [{"property": "video", "type": "video",
"data": "https://example.com/clip.mp4"}],
})
# 3. a collection over that bucket, running the extractor
collection = requests.post(f"{API}/v1/collections", headers=H, json={
"collection_name": "video-to-chapters",
"source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
"feature_extractor": {"feature_extractor_name": "multimodal_extractor", "version": "v1"},
}).json()
# 4. run extraction over the bucket
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
"collection_ids": [collection["collection_id"]],
"auto_submit": True,
})
# 5. read the output
docs = requests.get(
f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
).json()
print(docs)Use Cases
Supported Input Formats
Quick Info
Run it over a library
Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.
Frequently Asked Questions
Related Converters
Video to Summary
Produce concise written summaries of video content by combining transcript analysis, scene understanding, and key moment detection. Summaries can be formatted as paragraphs, bullet points, or structured chapters.
Video to Scenes
Automatically segment videos into individual scenes using visual and audio cue detection. Each scene includes a start and end timestamp, a representative keyframe, a descriptive label, and a confidence score for the detected boundary.
Video to Transcript
Extract a clean, time-stamped transcript from any video file with speaker diarization, punctuation restoration, and paragraph segmentation. Optimized for interviews, meetings, lectures, and multi-speaker content where accurate attribution matters.
Audio to Chapters
Automatically segment audio recordings into topic-based chapters with titles, timestamps, and summaries. Uses speech transcription combined with topic modeling to detect natural topic boundaries in podcasts, lectures, meetings, and audiobooks.
Ready to convert video to chapters?
Start using the Mixpeek Video to Chapters in minutes. Sign up for a free API key and follow the documentation to get started.