aimv2-large-patch14-native
by apple
Multimodal autoregressive vision encoder outperforming CLIP and SigLIP on understanding tasks
apple/aimv2-large-patch14-nativemixpeek://image_extractor@v1/apple_aimv2_large_v1Overview
AIMv2 Large is Apple's 309M-parameter vision encoder pre-trained with a multimodal autoregressive objective that pairs the encoder with a decoder autoregressively generating raw image patches and text tokens. Unlike contrastive models such as CLIP, AIMv2 captures fine-grained visual features through its generative pre-training, outperforming both CLIP and SigLIP on multimodal understanding benchmarks.
On Mixpeek, AIMv2 provides high-quality visual feature extraction for downstream tasks like classification, grounding, and retrieval. Its native resolution variant accepts variable-size images without resizing artifacts, making it particularly effective for document images, satellite imagery, and other content where resolution matters.
Architecture
Vision Transformer with 24 layers, 1024-dim hidden size, 8 attention heads, patch size 14. 309M parameters. Pre-trained with multimodal autoregressive objective using a paired text decoder. Native resolution variant supports variable input sizes without fixed resizing.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so AIMv2-large-patch14-native runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The vector name has to match a vector index on the collection.
vectors: { "image-embedding": yourVector },
payload: { source_key: "archive/2026/asset-00412" },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// image_extractor@v1 runs google/siglip-base-patch16-224
// (768-d) over a bucket, with no inference of your own.Capabilities
- Outperforms CLIP and SigLIP on multimodal understanding benchmarks
- Native resolution input without resizing artifacts
- 1024-dimensional feature representations
- Strong transfer to localization, grounding, and classification
- Outperforms DINOv2 on open-vocabulary detection
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| ImageNet-1k (frozen trunk, 3B variant) | Top-1 Accuracy | 89.5% | Fini et al., 2024: arXiv 2411.14402 |
| Multimodal understanding (avg) | Score | Outperforms CLIP ViT-L & SigLIP | Fini et al., 2024: arXiv 2411.14402 |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Multimodal Autoregressive Pre-training of Large Vision Encoders
arxiv.orgBuild a pipeline with aimv2-large-patch14-native
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free