Can AI Be a Participant in Social Experiments?
CWSF · 2026 Digital Technology
Overview
Social experiments are a part of human-centered social research. They are often costly and difficult to conduct. Some topics may even be unethical due to lasting harm. In this project, I explored whether AI could replace humans as an active participant in social experiments. I ran a survey in my school and also fed the same questionnaire to AI. I then compared human and AI responses. I conducted a thorough comparative analysis with the data sets, and the results show that, for topics involving applying knowledge and making predictions, AI can perform similar to humans and function not only as a mechanical tool but also as an intelligent social agent. My findings suggest that AI, as an active participant in some areas of human-centered research, can help save time and money and avoid ethical issues so that researchers can conduct more experiments and broaden their explorations.
Video
This video could not be played here. Watch it on the original project page.
Why?
I am interested in both AI and social science. I find it fascinating that some fields of social sciences, such as economics, psychology and political science, use lab experiments to study human behaviour and social outcomes. But social experiments are very expensive, time consuming, hard to find participants for, and sometimes even dangerous for human participants. Many social experiments do not end up happening because of these factors.
AI learns by analyzing massive amounts of data where algorithms identify patterns, make predictions, and adjust parameters to reduce errors. This process is called iterative training, which is mostly driven by machine learning and it mimics human learning to become more accurate over time with lots of practice.
How about we use AI as an active participant in social experiments? I hypothesize that AI responds similarly to humans so that it is able to participate in suitable experiments. If it works well, it will solve all of the problems I mentioned above. I believe that many can benefit from my project, for example, researchers in social sciences, businesses who need to understand customer spending behaviour, and insurance companies who need to accurately predict risk levels, etc.
To test this hypothesis, I used two popular Large Language Models (LLMs), ChatGPT and Google AI, to study how AI responds to social topics. I ran a survey in my school and the adjoining high school with the same questionnaires I fed to AI. I then conducted a comparative analysis on AI and human responses.
How?
For my project, I did two sets of survey with middle and high school students. The FLASF Questionnaire is the one I used for our regional science fair. The participants were Grade 8 students from my program at the Calvin Park Public School, Kingston, Ontario. The CWSF Questionnaire is an extended survey I did with high-schoolers at LCVI, which shares the same location as my school.
My questionnaires have three categories of questions: opinion, knowledge and decision. The first category is personality type of questions like favourite colours. The second category is knowledge-based questions, which ask about economic related terms. The third category is decision-making/behaviour-based questions, which include scenarios with two or more options.
After each survey, I gave the same questionnaire to AI and experimented with different prompts and varying amounts of background information to see if they help AI's answers become more accurate and relevant.
Finally, I compared the two sets of results from AI and the student groups to see how reliable AI is for being a participant in social experiments. I chose to use two AI models, ChatGPT and Google AI, to see if they are consistent with each other in generating responses.
All images, with the exception of the questionnaires in the HOW section, and the analytical graphs in the WHAT section, are generated using ChatGPT (OpenAI).
What?
In Figures 1-3, I selected a few questions from each category to report the comparison. The human group for these figures included 20 middle-schoolers. Figure 1 is for an opinion question asking about one's favourite colour. Google AI and ChatGPT both reported the most popular colour, blue, in a percentage very close to the student responses. Google AI is also close to the student group in the order and percentage of favourite colours, while ChatGPT has the order mostly consistent but percentages different from students.
Figure 2 is a knowledge question about inflation and unemployment. Both AI models answered close to the student group, although their answers about unemployment are quite different from the human response. I also noticed that in this example both AI models' answers are also very close to each other. This may be because inflation and unemployment are common economic concepts and so both models had a lot of information about them.
Figures 1-3 suggest that AI stayed decently accurate on the tested topics. However, AI struggled sometimes with decision/behaviour questions, such as those in Figure 3 about Taylor Swift and math. Students sometimes seemed to think outside the box and make creative answers. AI, however, is not as accurate in this category as it is for knowledge questions. Also, I noticed that AI's opinions were more easily influenced by those I prompted it. Instead, the students' responses were more stable and not easily influenced by outside opinions. There is a caveat to the results in Figure 3. I had a sample of only 20 students in the human group, who were from the same grade of the same program. Their opinions might have been biased due to similar backgrounds. For experiments, a group of 20 participants is a small sample size and could lead to results that are different from a much larger sample.
Figures 4-6 summarize my survey. I compared human and AI performances in five perspectives: accuracy, creativity, detail, personality, and consistency of one's opinion. To get the five measures, I assigned scores to each participant's answers. Accuracy is for knowledge questions, creativity for decision, detail for knowledge and decision, personality for opinion and decision, and consistency for opinion questions.
Figure 4 compares accuracy of answers to four knowledge questions I chose in the two questionnaires. Figure 5 reports the creativity gap between human and AI responses. Figure 6 compares human and AI performances in all five measures.
The AI vs human comparison from Figure 6 has interesting results: First, the student groups seem to be "better rounded" than AI models because their answers covered a wider range across all five measures. Second, AI models out-performed the student groups in terms of accuracy and detail, but lacked creativity, personality and consistency. Third, there are also differences in the responses from the middle-schoolers and the high-schoolers. Relative to high-schoolers, middle-schoolers' answers were stronger in personality but lacked detail. Fourth, Google AI and ChatGPT performed similarly to each other.
So What?
By working on this project, I experienced the challenges of conducting social experiments. It took a lot of time and effort planning, organizing, designing questionnaires, recruiting participants, giving instructions, and collecting questionnaires/permission forms back. I now understand even more the value of having AI as a participant in social experiments to reduce costs and relieve some human involvements.
Of course, my findings make it clear that AI still lacks human creativity, personality and consistency. But I think currently researchers can at least use AI to do "test runs". This is like before the actual launch of a spaceship, scientists use AI to make simulations. We can use AI as participants to get some preliminary predictions of a social experiment. This will already be helpful because it saves money and helps prevent problems. Moreover, for experiments could be harmful to humans, AI is the "next best thing" to human participants.
A direction to better fit AI for social experiments is to train it to pay more attention to personality/attitude/feelings, etc. Scientists have been working on using prompt engineering to help AI give more 'human like' responses. It is about finding ways to frame or clarify contexts so that AI can make more accurate and relevant outputs. There are also case studies about AI displaying human like emotions. These most recent research developments are so promising that I believe using AI as a participant in some social experiments will soon be reality.
What's Next?
To improve my project, an obvious step is to run the survey with much larger participant groups, which will allow me more reliable results from the human group and more accurate comparison for human vs AI. Another important step is to expand topics to test AI’s strengths and weaknesses regarding various topics. Finally, it would also be interesting to actually conduct an experiment rather than simply a survey so that I could actually observe the outcome of human decisions and see if AI participants led to a similar outcome.
Thanks
First of all, I am giving a shout-out to everyone who participated in my survey at the Calvin Park Public School and the Loyalist Collegiate and Vocational Institute (LCVI). Your volunteering provided valuable data for my project. I couldn't have done it without your support! I also would like to thank my science teacher, Mr. Candela, for the guidance on my project, and my homeroom teacher, Ms. Williams, for helping me arrange for surveys. I am grateful to Kamélia, Samuel, Sue, and the Best-of-Fair Judges of the Frontenac, Lennox and Addington Science Fair (FLASF) for providing me valuable feedback and helping me prepare for the Canada-Wide Science Fair. Last but not least, I am deeply grateful to my parents, Amy and Bryan, for their unwavering love and support for me to chase my dreams!!!
References
Ayad, R. (2023, August 31). Can we reverse engineer our social behaviour using AI? Faculty of Arts & Science. https://www.artsci.utoronto.ca/news/can-we-reverse-engineer-our-social-behaviour-using-ai
Barnow, B. S. (2010). Setting up social experiments: The Good, the Bad, and the Ugly. Zeitschrift Für ArbeitsmarktForschung, 43(2), 91–105. https://doi.org/10.1007/s12651-010-0042-6
https://www.facebook.com/verywell. (2019). 23 Great Psychology Experiment Ideas to Explore. Verywell Mind. https://www.verywellmind.com/psychology-experiment-ideas-2795669
Kendra Cherry. (2023, November 14). The Most Notorious Social Psychology Experiments. Verywell Mind. https://www.verywellmind.com/famous-social-psychology-experiments-2795667
Kimmel, A. (2023). 2 Ethical Issues in Social Influence Research Get access Arrow. Oup.com. https://academic.oup.com/edited-volume/28375/chapter-abstract/215259438?redirectedFrom=fulltext
Labvanced. (2024). 5 Famous & Classic Experiments | Social Psychology | Research. Labvanced.com. https://www.labvanced.com/content/research/en/blog/2024-04-5-famous-social-psychology-experiments/
Miller, K. (2025). Welcome To Zscaler Directory Authentication. Stanford.edu. https://news.stanford.edu/stories/2025/07/ai-social-science-research-simulated-human-subjects
Mishra, A. (2026). RISE Research. Riseglobaleducation.com. https://riseglobaleducation.com/blogs/social-psychology-experiments-you-can-replicate-at-school
Mojahar, A. (2025). Gemini vs Claude for Coding in 2025: We Tested Both. Index.dev. https://www.index.dev/blog/gemini-vs-claude-for-coding
ocufa_admin. (2023, August 22). Beyond the hype: How AI could change the game for social science research - Academic Matters. Academic Matters. https://academicmatters.ca/beyond-the-hype-how-ai-could-change-the-game-for-social-science-research/
Osamu Ekhator. (2025, April 23). I tested Gemini vs. Claude with 10 prompts: here’s the winner. Techpoint Africa. https://techpoint.africa/guide/i-tested-gemini-vs-claude-with-10-prompts/
Roberts, W., McKee, S. A., Miranda, R., & Barnett, N. P. (2024). Navigating ethical challenges in psychological research involving digital remote technologies and people who use alcohol or drugs. American Psychologist, 79(1), 24–38. https://doi.org/10.1037/amp0001193
Sanbonmatsu, D. M., Cooley, E. H., & Butner, J. E. (2021). The Impact of Complexity on Methods and Findings in Psychological Science. Frontiers in Psychology, 11(1), 580111. https://doi.org/10.3389/fpsyg.2020.580111
Scott, C. (2025, October 12). My Honest Take: Google AI Plus vs ChatGPT Plus. Medium. https://codyscottai.medium.com/my-honest-take-google-ai-plus-vs-chatgpt-plus-46a21630bc9a
Seeds, T. (2022, September 19). Social Science is Hard! Medium. https://medium.com/@theo.seeds/social-science-is-hard-ee5c6d67b63e
Sofroniew et al. (2026). Emotion Concepts and their Function in a Large Language Model. https://transformer-circuits.pub/2026/emotions/index.html
Zitter, L. (2025, February 28). Gemini vs. ChatGPT: What’s the difference? | TechTarget. Enterprise AI. https://www.techtarget.com/searchenterpriseai/tip/Gemini-vs-ChatGPT-Whats-the-difference
Images (21)
Awards (1)
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
CYSF · 2026
Can You Spot the Bot? Testing AI Detector Accuracy, Patterns and Use with Human, Humanized and AI Texts
ISEF · 2026
Characterizing Cognitive Dynamics in AI: Designing Virtual Worlds for a Comparative Analysis of Persona-Driven Behavior in Individual and Collaborative Agentic Systems
ISEF · 2023
Artificial Genius: Testing the Abilities of Artificial Intelligence in a More Turing Way
CWSF · 2026
The Impacts of Artificial Intelligence: Is It Positive or Negative
CYSF · 2025
AI vs. Human Authors: Who writes better stories?
ISEF · 2026
AI-Based System for Analyzing and Supporting Student Engagement in the Classroom
ISEF · 2025
AI Companion for ASD: Predicting Listener's Attention Using Multi-Modal Response Analysis
ISEF · 2022
Mental Health Risk Detection With Artificial Intelligence
Closest projects by meaning, across every fair and year in the corpus.