deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Blind test of 6,851 student votes puts Gemini ahead of ChatGPT and Claude on essays

StudyArena's blind-comparison data from August 2026 shows students chose Gemini's essay answers 39.6% of the time, ahead of Claude at 31.8% and OpenAI models at 29.2%.

Blind test of 6,851 student votes puts Gemini ahead of ChatGPT and Claude on essays

Gemini leads a 6,851-vote blind comparison

Students preferred Gemini's essay writing to ChatGPT's and Claude's when model names were hidden, according to new data from StudyArena, a site where students compare AI answers side by side and learn which model produced the winner only after voting. The findings, which reached Hacker News's front page on August 25, draw on 6,851 eligible blind votes from the platform's August 2026 dataset.

For writing and essay tasks, responses from Google's Gemini family were picked 39.6% of the time, ahead of Anthropic's Claude at 31.8% and OpenAI models, including ChatGPT, at 29.2%. StudyArena describes the dataset as de-identified aggregate production data, with internal and admin activity and ineligible ballots excluded, and current model variants grouped by provider family. The arena covers the latest flagship releases, namely Gemini 3.1 Pro, Claude Opus 5 and GPT-5.6 Sol, rather than the older generations still cited in many comparisons.

The blind format is the crux: a student who already pays for one model, or believes another one writes better, still has to judge the text on the page rather than the logo.

Longer answers tended to win

Length correlated with preference. The selected response was on average 37% longer than the alternatives it beat, and the longest answer won 47.7% of decisive writing ballots. Short answers were not hopeless, since the shortest response still won 25.0% of the time, but brevity was clearly not the winning strategy in this pool.

Higher reasoning settings scored worse

The more counterintuitive finding concerns reasoning effort. Choice rates fell as effort labels rose: low-effort settings won 40.7% of the time, medium 33.3%, default 30.8% and high 29.5%.

StudyArena's explanation is that extra reasoning gives a model room to pile on points, qualifications and repetition, which can help with a proof or a research plan but tends to weaken prose. Its practical advice for essay work is to start at a normal or low effort setting and move higher only when the difficult part is reasoning through evidence rather than writing the sentence.

Each provider led a different stage

Although Gemini took the overall writing recommendation, the specialty numbers split by task. Gemini led feedback on existing drafts at 41.7%, making it StudyArena's suggested starting point when a draft exists and needs diagnosis. Claude led assignment planning at 43.2%, particularly for generating competing structures and weighing their tradeoffs. OpenAI's models led research work at 39.3%, such as mapping claims, objections and facts that need verification.

In other words, which model came out ahead depended on which part of the essay process was being tested.

How much to trust the numbers

Some caution is warranted. The data comes from a single platform with a commercial interest in its own comparison tool, the voting population consists of self-selected students rather than a random sample, and no confidence intervals or error margins are published. StudyArena also discloses that it may use AI to help draft and analyze its articles. The leaderboard is live and will keep shifting as new votes arrive, so the ranking is a snapshot rather than a settled verdict.

Why it matters

Essay help is one of the most common ways students use chatbots, and a measurable preference shift at this scale is an early signal of where consumer loyalty may drift. The results also carry two practical takeaways beyond the ranking itself: leadership is task-dependent, with each provider strongest at a different stage of the writing process, and turning up reasoning effort can actively hurt prose output, a setting change users can test immediately. Finally, blind arenas are emerging as a useful complement to lab benchmarks, because they measure what people actually prefer in real tasks rather than what scores well on tests.

  • #gemini
  • #chatgpt
  • #claude
  • #ai-writing
  • #blind-testing

Related posts