ARXIV:2603.06982 · 3D SHAPE RETRIEVAL · SUBMITTED 19 MAR · 21:31 UTC · FRESHNESS STALE

VerifiedSource: PDF linkedPartialPaperPack: 3 of 4 citation fields filledMissingMissing fields: authorsPartialProof: unverified proof status

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

arXiv

Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

Blocked on Code›Score8.0Evidence unverified

Opportunity summary

Pain Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

Evidence 0 refs | 0 sources | 33% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval. Recent approaches typically rely on bridging the domain gap between 2D images and 3D shapes…

METHOD

Full abstract

Image-based shape retrieval (IBSR) aims to retrieve 3D models from a database given a query image, hence addressing a classical task in computer vision, computer graphics, and robotics. Recent approaches typically rely on bridging the domain gap between 2D images and 3D shapes based on the use of multi-view renderings as well as task-specific metric learning to embed shapes and images into a common latent space. In contrast, we address IBSR through large-scale multi-modal pretraining and show that explicit view-based supervision is not required. Inspired by pre-aligned image--point-cloud encoders from ULIP and OpenShape that have been used for tasks such as 3D shape classification, we propose the use of pre-aligned image and shape encoders for zero-shot and standard IBSR by embedding images and point clouds into a shared representation space and performing retrieval via similarity search over compact single-embedding shape descriptors. This formulation allows skipping view synthesis and naturally enables zero-shot and cross-domain retrieval without retraining on the target database. We evaluate pre-aligned encoders in both zero-shot and supervised IBSR settings and additionally introduce a multi-modal hard contrastive loss (HCL) to further increase retrieval performance. Our evaluation demonstrates state-of-the-art performance, outperforming related methods on $Acc_{Top1}$ and $Acc_{Top10}$ for shape retrieval across multiple datasets, with best results observed for OpenShape combined with Point-BERT. Furthermore, training on our proposed multi-modal HCL yields dataset-dependent gains in standard instance retrieval tasks on shape-centric data, underscoring the value of pretraining and hard contrastive learning for 3D shape retrieval. The code will be made available via the project website.

RESULT

ScienceToStartup currently rates this 8.0/10 on the public viability pass. In contrast, we address IBSR through large-scale multi-modal pretraining and show that explicit view-based supervision is not required.

WHY NOW

3D Shape Retrieval moved forward this cycle; last verified April 2026. Public score 8.0/10.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score8.0

PainImprove 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

Evidence0 refs | 0 sources | 33% coverage

Blockermissing authors

Analysis summary

Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

VerifiedSource: PDF linkedPartialPaperPack: 3 of 4 citation fields filledMissingMissing fields: authorsPartialProof: unverified proof status

Competitive landscape

Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

Segment

3D Shape Retrieval

Adoption evidence

No public code link in the paper record yet

Commercial read

8.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "2930e58a-8939-4e2a-9b0c-515a9840a564", "arxiv_id": "2603.06982", "canonical_route": "/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "endpoints": { "paper_pack": "/api/v1/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning/paper-pack", "build_passport": "/api/v1/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning", "normalized_query": "2603.06982", "route": "/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "paper_ref": "optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning#webpage", "url": "https://sciencetostartup.com/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "name": "Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning", "description": "Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning#scholarlyArticle", "headline": "Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning", "description": "Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.", "url": "https://sciencetostartup.com/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning", "sameAs": "https://arxiv.org/abs/2603.06982", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2603.06982" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-03-07T01:54:35.000Z", "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 8 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "3D Shape Retrieval" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "3D Shape Retrieval", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "Optimizing Multi-Modal Models for Image-Based Shape Retrieva", "item": "https://sciencetostartup.com/paper/optimizing-multi-modal-models-for-image-based-shape-retrieval-the-role-of-pre-alignment-and-hard-contrastive-learning" } ] } ] }

Competitive landscape

Improve 3D model retrieval from images using pre-trained multi-modal encoders and hard contrastive learning, enabling zero-shot and cross-domain retrieval.

Segment

3D Shape Retrieval

Adoption evidence

No public code link in the paper record yet

Commercial read

8.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline