LLM Wars: A Benchmark of Large Language Models for Pain Point Extraction from Software Feedback
CWSF · 2026 Digital Technology Bronze Medal
Overview
This study benchmarks ten large language models on their ability to extract and classify software user pain points from thousands of real-world feedback records across four platforms: G2, Capterra, App Stores, and Reddit. Unstructured feedback contains valuable signals such as frustrations and unmet needs, but it is too large and inconsistent to analyze manually. Each model was scored using a novel two-metric framework: F1 accuracy against a 900-record gold standard validated by 958 software developers over 70 days, and semantic agreement across all 12,000 records. Results reveal that model selection depends entirely on the use case: a system prioritizing precision needs a different model than one optimizing for scale, every model collapsed on informal conversational data, and the most expensive model is not worth the cost for most deployments. The evaluation framework is designed to generalize to any domain where humans express problems in natural language.
Awards (2)
- Bronze Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2025
Product Review Summarization and Chatbot Service Based on LangChain for Consumers
ISEF · 2025
PsychSPT: A Novel AI System for Mental Health Assessment Using Large Language Models (LLMs)
ISEF · 2026
LLMs Know When We Are Watching: A Lightweight Framework to Quantify Evaluation Awareness
ISEF · 2024
Bias in Large Language Models (LLMs); Paving the Way for an Equitable Artificial General Intelligence (AGI)
Closest projects by meaning, across every fair and year in the corpus.