DA3-LARGE-1.1
by depth-anything
Depth Anything V3 Large checkpoint for spatial scene retrieval
depth-anything/DA3-LARGE-1.1mixpeek://image_extractor@v1/depth_anything_v3_large_v1Overview
Depth Anything V3 Large is a newer monocular depth checkpoint for estimating dense depth maps from images. It improves the depth channel available to perception pipelines where spatial layout matters as much as semantic content.
On Mixpeek, Depth Anything V3 Large can enrich images or video frames with depth-derived metadata such as foreground ratio, depth bands, relative distance, and scene layout. Agents can use those signals to find clips by spatial evidence before calling a heavier reasoning model.
Architecture
Depth estimation model from the Depth Anything V3 family. It produces dense per-pixel depth maps from single images and is exposed through the HuggingFace depth-estimation pipeline.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so DA3-LARGE-1.1 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// Boxes, masks, depth maps and anomaly scores are structured
// results, not vectors. They go in payload and are reachable
// through pre_filters on a retriever, not through similarity.
payload: {
detections: modelOutput,
source_key: "archive/2026/asset-00412",
},
},
],
}),
},
);
// No managed alternative for an open label set. Two extractors do emit a
// bbox, for the one thing each detects: document_graph_extractor@v1 per
// layout block, face_identity_extractor@v1 per face. Nothing ships that
// returns masks, depth maps or anomaly scores.Capabilities
- Monocular dense depth estimation
- Spatial metadata for image and video retrieval
- Useful with segmentation and visual embeddings
- Apache 2.0 license
Use Cases on Mixpeek
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Depth Anything V3 Large 1.1
arxiv.orgBuild a pipeline with DA3-LARGE-1.1
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free