LLM Performance

CSEF · 2026 Mathematical Sciences (Junior Division)

Overview

We use AI in daily life, sometimes involuntarily (such as Google Search Overview). However, some AI models might be more reliable or structured than others. In this project, we tested six different text-generating AI models (LLMs) in order to find which of these LLMs is the most reliable. We tested these models by feeding them three uniquely themed prompts, and graded them according to how well they adhered to three grading criteria. We predicted that Claude AI would perform the best, which was incorrect. Claude incorrectly answered the mathematical question, averaging 81 out of 100 points. Google Gemini, however, boasted an average of 90.3 points out of 100, giving it first place. This, however, does not mean Gemini is the best LLM. What this does mean, however, is that Gemini met our standards the best. Further investigation could be performed to confirm these scores.

Competition history

  • CSEF 2026 Mathematical Sciences (Junior Division) · Entry J-14-13

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google