How SANet Slashes Multimodal AI Training Costs by 90% While Boosting Accuracy
We compile, generate and translate using Artificial Intelligence from the below given source. Macro Micro News is responsible for its editorial publication.
Global, Source: Global, Source
The landscape of Multimodal Large Language Models is shifting rapidly. Systems like LLaVA, InternVL, and Qwen-VL have transformed how machines interpret visual data. Yet a persistent hurdle remains. Creating high-quality instruction-tuning datasets demands significant resources. The process is costly and time-intensive. Often the resulting data contains noise or redundancy. Traditional scaling methods introduce weak alignment between images and text. This imbalance leads to inefficient training cycles.
Researchers have responded with SANet, a Selection–Augmentation Network. This framework offers a novel approach to efficient multimodal instruction data construction. It operates under fixed data budgets. Unlike static selection methods that simply filter existing data, SANet uses an iterative closed-loop process. It involves selection, refinement, supplementation, and re-evaluation. This cycle ensures only the most valuable samples contribute to model training. Computational overhead drops significantly while performance stays high.
SANet relies on two core mechanisms. Multi-Criteria Selection identifies candidate samples based on distribution coverage and optimization relevance. Bi-Level Augmentation refines individual samples and supplements under-represented capabilities. Crucially, augmented candidates do not enter the final subset immediately. They undergo subsequent re-selection stages. This extra step guarantees quality before inclusion.
Experiments using a fixed 6% data budget from the LLaVA-1.5 dataset highlight SANet's efficiency. The model achieved an average score of 61.88 across ScienceQA, MMBench, and TextVQA benchmarks. This result outperformed random selection and other baselines like LESS-MM. The compact subset reduced end-to-end training time to approximately 35 hours. Full-data fine-tuning required 67 hours. The constructed subset also showed strong cross-model reusability. It performed effectively when applied to different backbones such as Qwen2-VL-7B and InternVL2.5-8B. No backbone-specific reconstruction was needed.
Analysis reveals some limitations in fine-grained compositional reasoning tasks. Performance on the GQA benchmark shows room for improvement. The augmentation process occasionally introduces subtle reasoning or visual-grounding errors. These affect performance on highly specific evaluation metrics. Despite this, SANet marks a significant advance in data-centric AI. It offers a practical pipeline for training next-generation multimodal models. By decoupling data construction from model training, organizations can build reusable datasets. This accelerates development cycles and reduces reliance on massive computational resources.