ARXIV:2603.11873 · LLM OPTIMIZATION · SUBMITTED 02 APR · 02:30 UTC · FRESHNESS STALE

VerifiedSource: PDF linkedPartialPaperPack: 3 of 4 citation fields filledMissingMissing fields: authorsPartialProof: unverified proof status

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

arXiv

AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

Blocked on Code›Score3.0Evidence unverified

Opportunity summary

Pain AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

Evidence 0 refs | 0 sources | 17% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy. However, this architectural enhancement comes at a steep cost: despite minimal increases in computational load, the inference latency often skyrockets, leading…

METHOD

Full abstract

The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Models (LLMs). However, this architectural enhancement comes at a steep cost: despite minimal increases in computational load, the inference latency often skyrockets, leading to decoding speeds slowing by over 2.5 times. Through a fine-grained performance analysis, we pinpoint the primary bottleneck not in the computation itself, but in the severe overhead from fragmented, sequential CUDA kernel launches required for conventional dynamic routing. To address this challenge, we introduce AdaFuse, a framework built on a tight co-design between the algorithm and the underlying hardware system to enable efficient dynamic adapter execution. Departing from conventional layer-wise or block-wise routing, AdaFuse employs a token-level pre-gating strategy, which makes a single, global routing decision for all adapter layers before a token is processed. This "decide-once, apply-everywhere" approach effectively staticizes the execution path for each token, creating an opportunity for holistic optimization. We capitalize on this by developing a custom CUDA kernel that performs a fused switching operation, merging the parameters of all selected LoRA adapters into the backbone model in a single, efficient pass. Experimental results on popular open-source LLMs show that AdaFuse achieves accuracy on par with state-of-the-art dynamic adapters while drastically cutting decoding latency by a factor of over 2.4x, thereby bridging the gap between model capability and inference efficiency.

RESULT

ScienceToStartup currently rates this 3.0/10 on the public viability pass. To address this challenge, we introduce AdaFuse, a framework built on a tight co-design between the algorithm and the underlying hardware system to enable…

WHY NOW

LLM Optimization moved forward this cycle; last verified April 2026. Public score 3.0/10.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score3.0

PainAdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

Evidence0 refs | 0 sources | 17% coverage

Blockermissing authors

Analysis summary

AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

VerifiedSource: PDF linkedPartialPaperPack: 3 of 4 citation fields filledMissingMissing fields: authorsPartialProof: unverified proof status

Competitive landscape

AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

Segment

LLM Optimization

Adoption evidence

No public code link in the paper record yet

Commercial read

3.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "7a266cc2-0015-4ad1-a810-5ae6f66b0023", "arxiv_id": "2603.11873", "canonical_route": "/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "endpoints": { "paper_pack": "/api/v1/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization/paper-pack", "build_passport": "/api/v1/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization", "normalized_query": "2603.11873", "route": "/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "paper_ref": "adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization#webpage", "url": "https://sciencetostartup.com/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "name": "AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization", "description": "AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization#scholarlyArticle", "headline": "AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization", "description": "AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.", "url": "https://sciencetostartup.com/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization", "sameAs": "https://arxiv.org/abs/2603.11873", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2603.11873" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-03-12T12:46:42.000Z", "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 3 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "LLM Optimization" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "LLM Optimization", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "AdaFuse: Accelerating Dynamic Adapter Inference via Token-Le", "item": "https://sciencetostartup.com/paper/adafuse-accelerating-dynamic-adapter-inference-via-token-level-pre-gating-and-fused-kernel-optimization" } ] } ] }

Competitive landscape

AdaFuse optimizes dynamic adapter inference for LLMs, significantly reducing latency while maintaining accuracy.

Segment

LLM Optimization

Adoption evidence

No public code link in the paper record yet

Commercial read

3.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Related Resources

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline