Train an Open-Source Large Language Model to Accurately Solve Math Problems
AJAS · 2025 Robotics and Intelligent Machines (inferred)
Overview
Most Large Language Models (LLMs) are trained to answer general questions, but these models often fall short when asked to solve complex mathematical problems. Though GPT-4 vision is considered to be one of the best LLMs for general use, it is not an open-source model and is very expensive to use at large scales. Even the leading open-source LLMs encounter issues that result in incorrect answers when presented with SAT-level math questions. There is a growing demand for open-source models that can solve these math questions accurately to help students with SAT preparation, and the goal of this research was to address this demand by using fine-tuning methods to enhance the accuracy of an open-source LLM by at least 10%. The LoRA fine-tuning approach was used to finetune WizardMath, an open-source model. It was trained on one type of SAT questions which was regenerated to provide a dataset with different numbers. The pre-trained model answered 0% of the questions correctly. The trained model answered 80% of the questions correctly. The fine-tuned model's performance significantly improved. This research will benefit a wide range of people, from mathematicians to researchers working on fine-tuning with LoRA.
Competition history
- AJAS 2025
Related projects
CSEF · 2026
Multi-Label LLM Pretraining With A Smaller Teacher Model
CSEF · 2016
Using Artificial Intelligence Systems for Autonomous Visual Comprehension and Handwriting Generation
ISEF · 2026
MERIT: Mechanistic Explainability of Reasoning Integrity and Transparency
ISEF · 2016
Using Artificial Intelligence Systems for Autonomous Visual Comprehension and Handwriting Generation
ISEF · 2025
OATNet: A Computational and Mathematical Model of a Novel Neural Network Architecture Utilizing Ternary Weight Decomposition and Element-Wise Methods for Mitigating Computational Complexity
CWSF · 2026
LLM Wars: A Benchmark of Large Language Models for Pain Point Extraction from Software Feedback
CSEF · 2014
A Machine Learning Model for Automated Semantic Short Essay Assessment through Random Forest Based Ensembles and NLP
ISEF · 2024
Voicemath: A Calculator to Convert Commonly Spoken Mathematical Speech Into Equations Using Natural Language Processing
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: AAAS Annual Meeting (Confex) / American Junior Academy of Science