ARXIV:2605.14773 · LLM TRAINING OPTIMIZATION · SUBMITTED 15 MAY · 20:12 UTC · FRESHNESS FRESH

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training

Suorong Yang · Hanqi Zhu · Hai Gan · Fangjian Su · Guang Li · Furao Shen · +1 at arXiv

A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

Ship in 2-4 weeks›Score7.0Evidence unverified

Opportunity summary

Pain A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

Evidence 0 refs | 0 sources | 0% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models. However, existing methods mainly focus on designing sample-importance criteria, i.e., deciding what to select, while typically…

METHOD

Full abstract

Data selection accelerates training by identifying representative training data while preserving model performance. However, existing methods mainly focus on designing sample-importance criteria, i.e., deciding what to select, while typically fixing the selected data volume as the target ratio throughout training. Thus, they are often dynamic in sample identity but static in data volume. In this work, we revisit data selection from an optimization perspective and show that selected-data training induces an implicit regularization effect modulated by the instantaneous selection ratio. This reveals a key trade-off: lower ratios amplify selection-induced regularization, whereas higher ratios preserve data coverage and optimization fidelity. Motivated by this insight, we propose PODS, a Plug-and-play Oscillatory Data-volume Scheduling framework. Rather than introducing another sample-scoring metric, PODS serves as a lightweight module that dynamically schedules how much data to select over training. Under the target selection ratio, PODS alternates between low-ratio regularization phases and high-ratio recovery phases to exploit selection-induced regularization without sacrificing optimization stability. With its lightweight, ratio-level, and task-agnostic design, PODS is compatible with existing static and dynamic selection methods and broadly applicable across training paradigms. Experiments across various datasets, architectures, and tasks show that PODS consistently improves the efficiency-generalization trade-off, e.g., reducing ImageNet-1k training cost by 50% with improved accuracy and accelerating LLM instruction tuning by over 2x without performance degradation.

RESULT

ScienceToStartup currently rates this 7.0/10 on the public viability pass. In this work, we revisit data selection from an optimization perspective and show that selected-data training induces an implicit regularization effect modulated by the…

WHY NOW

LLM Training Optimization moved forward this cycle; last verified May 2026. Public score 7.0/10. Production flags indicate code availability.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score7.0

PainA plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

Evidence0 refs | 0 sources | 0% coverage

Blockerno shell-level blocker reported

Analysis summary

A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

Competitive landscape

A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

Segment

LLM Training Optimization

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "a193ce86-1b90-4fe0-8f09-382c890d8d92", "arxiv_id": "2605.14773", "canonical_route": "/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "endpoints": { "paper_pack": "/api/v1/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training/paper-pack", "build_passport": "/api/v1/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training", "normalized_query": "2605.14773", "route": "/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "paper_ref": "beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training#webpage", "url": "https://sciencetostartup.com/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "name": "Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training", "description": "A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training#scholarlyArticle", "headline": "Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training", "description": "A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.", "url": "https://sciencetostartup.com/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training", "sameAs": "https://arxiv.org/abs/2605.14773", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2605.14773" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-05-14T12:37:11.000Z", "author": [ { "@type": "Person", "name": "Suorong Yang" }, { "@type": "Person", "name": "Hanqi Zhu" }, { "@type": "Person", "name": "Hai Gan" }, { "@type": "Person", "name": "Fangjian Su" }, { "@type": "Person", "name": "Guang Li" }, { "@type": "Person", "name": "Furao Shen" }, { "@type": "Person", "name": "Soujanya Poria" } ], "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 7 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "LLM Training Optimization" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "LLM Training Optimization", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "Beyond What to Select: A Plug-and-play Oscillatory Data-Volu", "item": "https://sciencetostartup.com/paper/beyond-what-to-select-a-plug-and-play-oscillatory-data-volume-scheduling-for-efficient-model-training" } ] } ] }

Competitive landscape

A plug-and-play framework that oscillates data volume during training to improve efficiency-generalization trade-offs for LLMs and other models.

Segment

LLM Training Optimization

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training

Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline