NEWVectors or files. Pick a path.Start →
    Models/microsoft/table-transformer-structure-recognition
    MIT

    table-transformer-structure-recognition

    by microsoft

    Rows, columns, headers and spanning cells from a table image, the step after finding the table

    Identifiers
    Model ID
    microsoft/table-transformer-structure-recognition
    Feature URI

    Deploy table-transformer-structure-recognition

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Finding a table on a page and reading its structure are two separate jobs, and this checkpoint does the second. Given an image of a table, it returns boxes for six things: the table, its columns, its rows, column headers, projected row headers and spanning cells. Those boxes turn a picture of a grid into cells you can pair with text, which is the step a PDF pipeline needs before a figure in row 4, column 3 becomes a searchable fact.

    It is a DETR detection transformer with 28.8 million parameters and 125 object queries, trained on PubTables-1M and released with the paper that introduced that dataset. It expects a cropped table, so it runs after microsoft/table-transformer-detection has located one. It is downloaded 1.1 million times a month.

    Its output is geometry. Cell text still comes from the PDF text layer or from OCR, and the rebuilt table has to be written out as text, for example one line per row with the column headers attached, before an embedding model can index it.

    Architecture

    Table Transformer: the DETR architecture with a ResNet-18 backbone and layer normalization applied before attention (the "normalize before" setting, per the card). d_model 256, 6 encoder and 6 decoder layers, 125 object queries, 28,847,819 parameters. Six labels: table, table column, table row, table column header, table projected row header, table spanning cell. The preprocessor resizes the shorter side to 800 pixels with a 1000-pixel maximum and normalizes with ImageNet mean and standard deviation.

    Mixpeek SDK Integration

    // Runs on your side (or on a single-tenant deployment with uploaded weights).
    // The model returns boxes; intersect row and column boxes to get cells, fill
    // them from the PDF text layer or OCR, then write each row out as text.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "filing-2026-q2-p14-table-1-row-4",
              // Embedding of the serialized row, from a text embedding model.
              vectors: { "text-embedding": yourRowVector },
              payload: {
                source_key: "filings/2026-q2.pdf",
                page: 14,
                row_text: "Segment: Cloud | Revenue: 4,210 | Operating margin: 18.2%",
              },
            },
          ],
        }),
      },
    );

    Capabilities

    • Six structure classes, including spanning cells and projected row headers
    • DETR set prediction with 125 queries per image, so no anchor boxes to configure
    • Pairs with the table detection checkpoint from the same repository
    • MIT license

    Use Cases on Mixpeek

    Rebuilding tables from scanned financial statements before indexing their figures
    Turning tables in research PDFs into rows an embedding model can read
    Flagging tables with merged or spanning cells for review
    Storing each cell value with its column header as a filterable field

    Performance

    Input SizeCropped table image, shorter side resized to 800 px (1000 px maximum)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The Hugging Face card publishes no accuracy or latency figures, and we have not measured it. The PubTables-1M paper reports structure recognition results for the models it trained.

    Frequently Asked Questions

    What is the difference between table-transformer-detection and table-transformer-structure-recognition?

    Detection takes a whole page and returns where the tables are. Structure recognition takes one cropped table and returns its columns, rows, headers and spanning cells. A pipeline runs them in that order.

    Does it read the text inside the cells?

    No. It returns boxes and labels. Intersect the row and column boxes to get cell regions, then take each cell's text from the PDF text layer when the file has one, or from OCR on a scan.

    How should a rebuilt table be indexed for search?

    Write each row as a line of text with its column headers attached, embed that line with a text embedding model, and keep the individual values in the document payload. The embedding makes a question like "cloud operating margin" find the row, and the payload fields make exact filters on the numbers possible.

    Does Mixpeek run this model for me?

    Not on the managed tier. No Mixpeek extractor loads these weights, so you run it on your side and upsert what the pipeline produces, or upload the weights on a single-tenant Enterprise deployment and have a custom plugin run them.

    Specification

    Organizationmicrosoft
    Retriever-
    Parameters29M
    LicenseMIT
    Downloads/moN/A
    Likes229

    Research Paper

    PubTables-1M: Towards comprehensive table extraction from unstructured documents

    arxiv.org

    Build a pipeline with table-transformer-structure-recognition

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free