ARXIV:2606.03626 · MULTIMODAL AI · SUBMITTED 03 JUN · 20:32 UTC · FRESHNESS FRESH

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

Chao Wen · Jacqueline Staub · Adish Singla · arXiv

A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

Ship in 2-4 weeks›Score6.0Evidence unverified

Opportunity summary

Pain A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

Evidence 0 refs | 4 sources | 83% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics. However, most prior work focuses on visual programming for productivity; it remains unclear how well current…

METHOD

Full abstract

Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for productivity; it remains unclear how well current VLMs perform on education-oriented visual programming and what factors limit their performance. To bridge this gap, we introduce TurtleAI, a benchmark containing 823 tasks curated based on real-world visual programming tasks in the Turtle Graphics domain. Solving these tasks requires models to perceive geometric patterns, reason about spatial relationships, and synthesize Python code that faithfully reproduces geometric patterns. We evaluate 20+ VLMs, including GPT-5, GPT-4o, and Qwen2-VL-72B, and find that they struggle significantly, with most achieving success rates below 30%. To address these limitations, we propose a data generation technique that requires only a small set of seed samples. Fine-tuning Qwen2-VL-72B on the resulting synthetic data yields an improvement of about 20% on real-world tasks. Our failure analysis reveals that GPT-4o struggles with spatial reasoning and precise visual replication, whereas fine-tuning primarily improves the alignment between visual reasoning and code implementation.

RESULT

ScienceToStartup currently rates this 6.0/10 on the public viability pass. Our failure analysis reveals that GPT-4o struggles with spatial reasoning and precise visual replication, whereas fine-tuning primarily improves the alignment between visual reasoning and…

WHY NOW

Multimodal AI moved forward this cycle; last verified June 2026. Public score 6.0/10. Implementation evidence is present through a linked repository.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score6.0

PainA benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

Evidence0 refs | 4 sources | 83% coverage

Blockerno shell-level blocker reported

Analysis summary

A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

Competitive landscape

A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

Segment

Multimodal AI

Adoption evidence

Public code linked for build inspection

Commercial read

6.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "17a0fdd8-a54f-49e5-baa0-c7374aecd82a", "arxiv_id": "2606.03626", "canonical_route": "/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "endpoints": { "paper_pack": "/api/v1/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics/paper-pack", "build_passport": "/api/v1/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics", "normalized_query": "2606.03626", "route": "/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "paper_ref": "turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics#webpage", "url": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "name": "TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics", "description": "A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics#scholarlyArticle", "headline": "TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics", "description": "A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.", "url": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics", "sameAs": "https://arxiv.org/abs/2606.03626", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2606.03626" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-06-02T13:25:05.000Z", "author": [ { "@type": "Person", "name": "Chao Wen" }, { "@type": "Person", "name": "Jacqueline Staub" }, { "@type": "Person", "name": "Adish Singla" } ], "codeRepository": "https://github.com/machine-teaching-group/acl2026-turtleai", "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 6 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "Multimodal AI" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code, repo url" } ] }, { "@type": "SoftwareSourceCode", "@id": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics#software", "name": "TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics - Source Code", "description": "A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.", "codeRepository": "https://github.com/machine-teaching-group/acl2026-turtleai", "url": "https://github.com/machine-teaching-group/acl2026-turtleai" }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "Multimodal AI", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "TurtleAI: Benchmarking Multimodal Models for Visual Programm", "item": "https://sciencetostartup.com/paper/turtleai-benchmarking-multimodal-models-for-visual-programming-in-turtle-graphics" } ] } ] }

Competitive landscape

A benchmark and fine-tuning approach for evaluating and improving multimodal models on visual programming tasks in Turtle Graphics.

Segment

Multimodal AI

Adoption evidence

Public code linked for build inspection

Commercial read

6.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline