codebert-base
by microsoft
Pre-trained model for code understanding and generation
microsoft/codebert-basemixpeek://document_extractor@v1/microsoft_codebert_base_v1Overview
CodeBERT is a bimodal pre-trained model for programming languages and natural language. It supports code search, code documentation generation, and code-to-code translation across 6 programming languages.
On Mixpeek, CodeBERT extracts and embeds code blocks from documents, enabling semantic search over code content, find code snippets by describing what they do in natural language.
Architecture
RoBERTa-base architecture (12 layers, 768-dim hidden, 12 attention heads) pre-trained on CodeSearchNet dataset with Masked Language Modeling and Replaced Token Detection on both NL and PL.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so codebert-base runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Natural language code search
- Code documentation generation
- 6 programming languages: Python, Java, JS, PHP, Ruby, Go
- 768-dimensional code embeddings
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| CodeSearchNet | MRR | 67.2 | Feng et al., 2020: Table 3 |
| Clone Detection (BigCloneBench) | F1 | 96.5% | Feng et al., 2020: Table 5 |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
arxiv.orgBuild a pipeline with codebert-base
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free