NEWVectors or files. Pick a path.Start →

    Feature Extractors

    After your data is connected, extractors run in parallel to pull out structured features, embeddings, entities, transcripts, and more.

    7 production extractors, each with a README, a live schema, and a Studio path

    Web Scraper + Multimodal Embeddings

    Crawl sites (docs, job boards, news, SPAs) and extract text, code & image embeddings in one pass.

    Text

    Multimodal Video/Audio/Image (Vertex v1 · Gemini v2)

    Unified embeddings for video, audio, image & text: FFmpeg scene/silence chunking, Whisper transcription, thumbnails.

    Multimodal

    Universal All-in-One (Gemini)

    One extractor for image, video, audio & documents: auto-detects modality and applies the right pipeline.

    Multimodal

    Multi-File Object Embeddings (Gemini)

    Embed ALL files of an object (images, PDFs, video, audio, text) into one 3072-D Gemini vector.

    Multimodal

    Document Layout Graph

    Decompose PDFs into spatial blocks: paragraphs, tables, forms, headers: with layout classification & confidence.

    Document

    Passthrough (Storage Only)

    Store and canonicalize objects with zero ML: metadata-only ingestion.

    Utility

    Scrolling/Marquee Text OCR

    Reads scrolling video text via phase-correlation band detection, panoramic stitching, and VLM OCR.

    Video

    What's new in extractors

    Full changelog