LLM Performance
CSEF · 2026 Mathematical Sciences (Junior Division)
Overview
We use AI in daily life, sometimes involuntarily (such as Google Search Overview). However, some AI models might be more reliable or structured than others. In this project, we tested six different text-generating AI models (LLMs) in order to find which of these LLMs is the most reliable. We tested these models by feeding them three uniquely themed prompts, and graded them according to how well they adhered to three grading criteria. We predicted that Claude AI would perform the best, which was incorrect. Claude incorrectly answered the mathematical question, averaging 81 out of 100 points. Google Gemini, however, boasted an average of 90.3 points out of 100, giving it first place. This, however, does not mean Gemini is the best LLM. What this does mean, however, is that Gemini met our standards the best. Further investigation could be performed to confirm these scores.
Competition history
- CSEF 2026
Related projects
CSEF · 2026
Does AI Have Personality?
CSEF · 2026
Trusting Minds vs. Machines
CSEF · 2026
Quantifying Variability in AI Outputs
CSEF · 2026
AI Propaganda: Who Can Tell?
ISEF · 2026
Promptly Speaking: The Relationship Between AI Prompts, the Results You Receive, & Their Effect on Student Learning
CSEF · 2026
Detection of AI-Generated Music Using Convolutional Neural Networks
CSEF · 2023
Exploring Academic Setting Applications of a Turing Test Derivative
ISEF · 2023
Artificial Genius: Testing the Abilities of Artificial Intelligence in a More Turing Way
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects