NEWVectors or files. Pick a path.Start →

    AI Model Hub

    Browse AI models for multimodal decomposition and recomposition pipelines: plug any model into your extractors.

    17,271 models available

    Showing 361-384 of 17,271 models

    Sentence Similarity

    Qdrant/bm25

    1.1M
    35
    transformers
    Image Text To Text

    unsloth/Muse-Glimmer-30B-GGUF

    1.1M
    519
    transformers
    Text Classification

    pysentimiento/robertuito-sentiment-analysis

    1.0M
    102
    pysentimiento
    Audio Classification

    alefiury/wav2vec2-large-xlsr-53-gender-recognition-librispeech

    1.0M
    48
    transformers
    Text Generation

    LiquidAI/LFM2.5-2.6B-GGUF

    1.0M
    327
    gguf
    Text Generation

    deepseek-ai/DeepSeek-V3-0324

    1.0M
    3,167
    transformers
    Text Generation

    deepseek-ai/DeepSeek-V3

    1.0M
    4,180
    transformers
    Image Text To Text

    RedHatAI/Qwen3.6-35B-A3B-NVFP4

    1.0M
    172
    transformers
    Image Text To Text

    LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V9-GGUF

    1.0M
    542
    hermes
    Image Text To Text

    shawnw3i/Huihui-Qwen3.6-27B-abliterated-AWQ-MTP

    1.0M
    13
    transformers
    Audio Classification

    audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim

    1.0M
    173
    transformers
    Sentence Similarity

    shibing624/text2vec-base-chinese

    1.0M
    801
    sentence-transformers
    Text Generation

    Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8

    1.0M
    200
    transformers
    Automatic Speech Recognition

    kresnik/wav2vec2-large-xlsr-korean

    1.0M
    56
    transformers
    Sentence Similarity

    Qdrant/bge-small-en-v1.5-onnx-Q

    1.0M
    2
    transformers
    Sentence Similarity

    sentence-transformers/distiluse-base-multilingual-cased-v1

    1.0M
    134
    sentence-transformers
    Feature Extraction

    unslothai/1

    1.0M
    1
    transformers
    Image Text To Text

    Qwen/Qwen3.5-2B-Base

    1.0M
    89
    transformers
    Image Text To Text

    LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V11-GGUF

    1.0M
    579
    hermes
    Image Segmentation

    ZhengPeng7/BiRefNet

    998K
    633
    birefnet
    Text Generation

    zai-org/GLM-5.2

    997K
    5,076
    transformers
    Text Generation

    OBLITERATUS/Qwen3.8-27B-OBLITERATED

    995K
    1,110
    mlx
    Text Generation

    deepseek-ai/DeepSeek-R1-0528-Qwen3-8B

    991K
    1,084
    transformers
    Text Generation

    legraphista/glm-4-9b-chat-IMat-GGUF

    990K
    5
    gguf
    16 / 720

    Choosing a model for multimodal retrieval

    Which of these models does Mixpeek actually run?

    Nine, and they are not the same thing as the catalog. The managed extractors run intfloat/multilingual-e5-large-instruct for text (1024-d), google/siglip-base-patch16-224 for images (768-d), google/vertex-multimodal (1408-d) and google/gemini-embedding-2 (3072-d) for unified multimodal, insightface ArcFace for faces (512-d), CLAP for audio fingerprints (512-d), facebook/dinov2-base for visual similarity and jinaai/jina-embeddings-v2-base-code for code (both 768-d, inside the web scraper). The rerank retriever stage runs BAAI/bge-reranker-v2-m3. Everything else in this catalog is documented here, not hosted here.

    Can I run any model from this catalog on Mixpeek?

    Not by naming it. No extractor takes a Hugging Face model id as a parameter, so there is no field to put one in. Three paths do work. Use a managed extractor and get the model it runs. Run the model on your own hardware and upsert the vectors through POST /v1/namespaces/{namespace_id}/documents/upsert, which stores them beside everything else. Or on Enterprise, upload the weights through POST /v1/namespaces/{namespace_id}/models, which accepts the huggingface format, and load them from a custom plugin.

    Should I use a text embedding model or a multimodal one?

    Ask whether a text query has to reach a non-text asset directly. If your video is searchable through its transcript and your images through their captions, a text model over that generated text is cheaper and usually more accurate, because retrieval quality on words is a solved problem and cross-modal alignment is not. If the query is 'find the shot that looks like this' or the visual content carries meaning no caption records, you need a shared space and a multimodal model. Most production systems run both indexes rather than choosing.

    What does the embedding dimension cost to store?

    A float32 vector is 4 bytes per dimension, so a million items costs 4 GB at 1024 dimensions, 3 GB at 768, and 12 GB at 3072. That is before any quantization and before payload. Models trained with Matryoshka representation learning, such as Gemini Embedding 2, let you truncate to a shorter prefix without re-encoding, so the width becomes an index-time decision rather than a model-selection one. Dimensions are fixed at namespace creation in Mixpeek, so changing a width later means re-indexing.

    Why does the download count on a model page differ from HuggingFace?

    It is the monthly figure from the HuggingFace API at the time of the last sync, not a live read, so it lags. The sync date is on each model page. Download count is a popularity signal and a poor quality signal: the most-downloaded model in a category is frequently an older checkpoint that a tutorial pinned years ago.

    All models

    Every model page in one place, 375 in total.