ARXIV:2603.26362 · VISION-LANGUAGE MODELS · SUBMITTED 30 MAR · 22:21 UTC · FRESHNESS STALE

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

MD Khalequzzaman Chowdhury Sayem · Mubarrat Tajoar Chowdhury · Yihalem Yimolal Tiruneh · Muneeb A. Khan · Muhammad Salman Ali · Binod Bhattarai · +1 at arXiv

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Ship in 2-4 weeks›Score7.0Evidence unverified

Opportunity summary

Pain A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Evidence 113 refs | 3 sources | 50% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLMs)…

METHOD

Full abstract

Understanding the fine-grained articulation of human hands is critical in high-stakes settings such as robot-assisted surgery, chip manufacturing, and AR/VR-based human-AI interaction. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLMs) struggle with fine-grained spatial reasoning, especially in interpreting complex and articulated hand poses. We introduce HandVQA, a large-scale diagnostic benchmark designed to evaluate VLMs' understanding of detailed hand anatomy through visual question answering. Built upon high-quality 3D hand datasets (FreiHAND, InterHand2.6M, FPHA), our benchmark includes over 1.6M controlled multiple-choice questions that probe spatial relationships between hand joints, such as angles, distances, and relative positions. We evaluate several state-of-the-art VLMs (LLaVA, DeepSeek and Qwen-VL) in both base and fine-tuned settings, using lightweight fine-tuning via LoRA. Our findings reveal systematic limitations in current models, including hallucinated finger parts, incorrect geometric interpretations, and poor generalization. HandVQA not only exposes these critical reasoning gaps but provides a validated path to improvement. We demonstrate that the 3D-grounded spatial knowledge learned from our benchmark transfers in a zero-shot setting, significantly improving accuracy of model on novel downstream tasks like hand gesture recognition (+10.33%) and hand-object interaction (+2.63%).

RESULT

ScienceToStartup currently rates this 7.0/10 on the public viability pass. We demonstrate that the 3D-grounded spatial knowledge learned from our benchmark transfers in a zero-shot setting, significantly improving accuracy of model on novel downstream…

WHY NOW

Vision-Language Models moved forward this cycle; last verified April 2026. Public score 7.0/10. Production flags indicate code availability.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score7.0

PainA diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Evidence113 refs | 3 sources | 50% coverage

Blockerno shell-level blocker reported

Analysis summary

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

MD Khalequzzaman Chowdhury Sayem · Mubarrat Tajoar Chowdhury · Yihalem Yimolal Tiruneh · Muneeb A. Khan · Muhammad Salman Ali · Binod Bhattarai · +1 at arXiv

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Competitive landscape

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Segment

Vision-Language Models

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "3d1ad394-fd6d-44da-926d-88df09bc58e0", "arxiv_id": "2603.26362", "canonical_route": "/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "endpoints": { "paper_pack": "/api/v1/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models/paper-pack", "build_passport": "/api/v1/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models", "normalized_query": "2603.26362", "route": "/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "paper_ref": "handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models#webpage", "url": "https://sciencetostartup.com/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "name": "HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models", "description": "A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models#scholarlyArticle", "headline": "HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models", "description": "A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.", "url": "https://sciencetostartup.com/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models", "sameAs": "https://arxiv.org/abs/2603.26362", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2603.26362" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-03-27T12:42:26.000Z", "author": [ { "@type": "Person", "name": "MD Khalequzzaman Chowdhury Sayem" }, { "@type": "Person", "name": "Mubarrat Tajoar Chowdhury" }, { "@type": "Person", "name": "Yihalem Yimolal Tiruneh" }, { "@type": "Person", "name": "Muneeb A. Khan" }, { "@type": "Person", "name": "Muhammad Salman Ali" }, { "@type": "Person", "name": "Binod Bhattarai" }, { "@type": "Person", "name": "Seungryul Baek" } ], "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 7 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "Vision-Language Models" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "Vision-Language Models", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "HandVQA: Diagnosing and Improving Fine-Grained Spatial Reaso", "item": "https://sciencetostartup.com/paper/handvqa-diagnosing-and-improving-fine-grained-spatial-reasoning-about-hands-in-vision-language-models" } ] } ] }

Competitive landscape

A diagnostic benchmark and fine-tuning method to significantly improve vision-language models' spatial reasoning about human hands, enabling applications in robotics and AR/VR.

Segment

Vision-Language Models

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline