Composed Image Retrieval With No Vision Model: 44% R@1 Zero-Shot
Summary
Composed image retrieval handles queries like 'this jacket, but in red, with a hood.' A University of Tokyo team does it with images represented purely as text: a caption plus an explicit attribute checklist at ingest, an LLM to merge the modification, plain text retrieval, attribute scoring, and an LLM reranker to verify the change happened. Zero-shot: 44.04% R@1 on CIRR, +8.79 over prior methods. Paper: arxiv.org/abs/2607.12621
About this video
Composed image retrieval handles queries like 'this jacket, but in red, with a hood.' A University of Tokyo team does it with images represented purely as text: a caption plus an explicit attribute checklist at ingest, an LLM to merge the modification, plain text retrieval, attribute scoring, and an LLM reranker to verify the change happened. Zero-shot: 44.04% R@1 on CIRR, +8.79 over prior methods. Paper: arxiv.org/abs/2607.12621