Macro Micro News Global Pulse. Local Truth.

How SANet Slashes Multimodal AI Training Costs by 90% While Boosting Accuracy

28 September 2026 · 2 min read

We compile, generate and translate using Artificial Intelligence from the below given source. Macro Micro News is responsible for its editorial publication.

Article image by Adrien
Image by Adrien

Global, Source: Global, Source

The landscape of Multimodal Large Language Models is shifting rapidly. Systems like LLaVA, InternVL, and Qwen-VL have transformed how machines interpret visual data. Yet a persistent hurdle remains. Creating high-quality instruction-tuning datasets demands significant resources. The process is costly and time-intensive. Often the resulting data contains noise or redundancy. Traditional scaling methods introduce weak alignment between images and text. This imbalance leads to inefficient training cycles.

Researchers have responded with SANet, a Selection–Augmentation Network. This framework offers a novel approach to efficient multimodal instruction data construction. It operates under fixed data budgets. Unlike static selection methods that simply filter existing data, SANet uses an iterative closed-loop process. It involves selection, refinement, supplementation, and re-evaluation. This cycle ensures only the most valuable samples contribute to model training. Computational overhead drops significantly while performance stays high.

SANet relies on two core mechanisms. Multi-Criteria Selection identifies candidate samples based on distribution coverage and optimization relevance. Bi-Level Augmentation refines individual samples and supplements under-represented capabilities. Crucially, augmented candidates do not enter the final subset immediately. They undergo subsequent re-selection stages. This extra step guarantees quality before inclusion.

Experiments using a fixed 6% data budget from the LLaVA-1.5 dataset highlight SANet's efficiency. The model achieved an average score of 61.88 across ScienceQA, MMBench, and TextVQA benchmarks. This result outperformed random selection and other baselines like LESS-MM. The compact subset reduced end-to-end training time to approximately 35 hours. Full-data fine-tuning required 67 hours. The constructed subset also showed strong cross-model reusability. It performed effectively when applied to different backbones such as Qwen2-VL-7B and InternVL2.5-8B. No backbone-specific reconstruction was needed.

Analysis reveals some limitations in fine-grained compositional reasoning tasks. Performance on the GQA benchmark shows room for improvement. The augmentation process occasionally introduces subtle reasoning or visual-grounding errors. These affect performance on highly specific evaluation metrics. Despite this, SANet marks a significant advance in data-centric AI. It offers a practical pipeline for training next-generation multimodal models. By decoupling data construction from model training, organizations can build reusable datasets. This accelerates development cycles and reduces reliance on massive computational resources.