NEWVectors or files. Pick a path.Start →
    Back to DiagramsBenchmarks

    Gemini Flash Metadata Extraction: The Headline Metric Picks the Wrong Model

    Seven Gemini Flash generations, one prompt, 143 real product photos. The F2 spread is 0.035, so the headline ranking is noise. The model that names the exact product type most often is fifth in it.

    Two rankings of seven Gemini Flash generations on metadata extraction over 143 product photos, close to inverted. Ranked by F2: 3.7 Flash 0.241, 3.1 Flash Lite 0.229, 3.6 Flash 0.228, 2.5 Flash 0.223, 3.5 Flash Lite 0.214, 3.5 Flash 0.209, 3 Flash preview 0.207, a spread of 0.035. Ranked by how often the model names the exact product type: 3.5 Flash Lite 61 percent, 3.1 Flash Lite 57, 3.6 Flash 50, 2.5 Flash 44, 3 Flash preview 44, 3.7 Flash 42, 3.5 Flash 41, a spread of 20 points. Gemini 3.5 Flash Lite sits fifth on F2 and first on product type, at 1502ms against 3.7 Flash's 3063ms.
    Seven Gemini Flash generations, one prompt, 143 real product photos. The F2 spread is 0.035, so the headline ranking is noise. The model that names the exact product type most often is fifth in it.

    Seven generations of Gemini Flash extracted metadata from the same 143 real e-commerce product photos with an identical prompt, so the model is the only variable. Ground truth is each merchant's own product title, and predicted terms are compared against it as sets.

    On F2, the metric the benchmark leads with, the whole field lands inside 0.035. Gemini 3.7 Flash is nominally first at 0.241 and the 3 Flash preview last at 0.207. A gap that small is not a result you should pick a model on.

    The second ranking is the one that decides whether a catalogue is searchable. How often does the model name the exact product type a shopper would type? There the field spreads 20 points, and the order is close to inverted. Gemini 3.5 Flash Lite names it 61 percent of the time against 3.7 Flash's 42 percent, and it does that in 1502ms against 3063ms.

    The mechanism is not mysterious. A lite model writes a shorter, cleaner tag list, so it has fewer chances to emit something wrong and its precision climbs. A terse list also lands on the obvious noun more reliably. If your pipeline needs the right product type on every asset and you are indexing at volume, that combination beats a marginally better F2.

    F2 is still the right primary metric for this task, which is why it leads. Recall and precision are not symmetric when the output is tags: a term the model never emits makes the asset unfindable, while a spurious term only adds noise to ranking. F2 weights recall four times as heavily to match that. One caveat belongs with the absolute numbers: F-scores compress when the reference term set is small, and a merchant title carries only a handful of scoreable terms, so recall and product-type hit rate carry more signal than the headline number.

    Gemini 3.5 Flash recorded two API failures during the run. That is in its row rather than retried away, because a benchmark that hides its failures is not one.

    The lesson generalises past Gemini. When a leaderboard's spread is smaller than the difference between what it measures and what your application needs, the leaderboard is selecting on noise. Rank by the axis that decides your outcome, and check whether the two orders agree before you trust either.

    Run this on your own data

    Mixpeek turns video, images, audio, and documents in your object storage into searchable, timestamped results through one API.

    Search your own data