OmniParser-v2.0
by microsoft
Screen parser that turns screenshots into structured UI elements for agents
microsoft/OmniParser-v2.0mixpeek://image_extractor@v1/microsoft_omniparser_v2_v1Overview
OmniParser v2 is Microsoft's screen parsing model for computer-use agents. It converts screenshots into structured elements by detecting interactable regions and captioning icons, so an LLM can reason over a screen as objects with coordinates and functions.
On Mixpeek, OmniParser is relevant for indexing UI recordings, app screenshots, support sessions, and agent traces. It makes visual interfaces searchable by element semantics instead of raw pixels alone.
Architecture
Two-model screen parser combining a fine-tuned YOLOv8 icon detector with a fine-tuned Florence-2 icon captioner. V2 adds cleaner icon grounding data and lower latency than V1.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so OmniParser-v2.0 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "multimodal-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Detects clickable and actionable UI regions
- Captions icons with functional semantics
- Converts screenshots into structured screen elements
- Useful with computer-use agents and GUI automation
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| ScreenSpot Pro | Average accuracy | 39.6 | Microsoft OmniParser v2 model card |
Performance
Best used for UI screenshots rather than natural scene imagery
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
OmniParser for Pure Vision Based GUI Agent
arxiv.orgBuild a pipeline with OmniParser-v2.0
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free