Rate My Face ChatGPT: AI Face Scores Run 1.7 Points Too Generous
Just another WordPress site
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
NEW YORK – September 29, 2026 – A test of whether people should rate my face ChatGPT style prompts found that large language models rate faces an average of 1.7 points higher than human raters do, and that no model tested returned a score below 5.5 out of 10.
looksmaxxing.guide ran the same ten photographs through ChatGPT (GPT-4o), Claude and Gemini, asking each for a numerical attractiveness score, the strongest and weakest features, and specific improvements. The same photographs went to a small human panel and to QOVES Studio, a dedicated facial-analysis tool.
The models averaged 7.1 out of 10. The human panel averaged 5.4 for the same faces. The dedicated tools averaged 5.8. The lowest score any model gave was 5.5. The human panel went as low as 3.
The reason is built into how the models are trained. Rating a face 4 out of 10 is harmful by most content-policy definitions, so a model designed to be helpful and harmless is pushed toward a diplomatic answer. The number is a compliment rather than a measurement.
Photo quality moved the score more than the face did. The same person shot in good and bad lighting drew scores differing by up to 1.8 points from the same model. The dedicated tool varied by 0.4 to 0.6 points across the same comparison.
The three models also agreed with each other closely, clustering within half a point on most photographs, which suggests they draw on the same training material rather than assessing independently.
Improvement advice was close to identical regardless of the face. Eight of the ten photographs came back with a version of the same suggestions: a skincare routine, a different hairstyle, hydration, posture.
The test also covered the “jailbreak” prompts circulating on TikTok, which ask a model to act as a ruthless casting agent and score honestly. A jailbroken 4 is no better calibrated than a default 7. There is no hidden real score behind the safety guidelines. Both the kindness and the harshness are generated by the prompt.
What the models did well was identify obvious features. When a model flagged strong jaw definition, clear skin, under-eye hollows or asymmetry, the human panel generally agreed. The models could not assess canthal tilt, midface ratio or the structural proportions that dedicated tools measure.
“People are asking a chatbot a question it is built not to answer honestly,” said Thomas Prommer, who publishes the site and works in AI and digital transformation. “A model trained to be helpful will not tell you that you are a four. That is not a flaw in the model, it is the model working as intended, and it means the number you get back is not a measurement of anything.”
The site’s recommendation is to use language models for what they are good at, which is style, grooming and skincare direction, and to disregard numerical scores from any AI source. Its final line is that five honest friends remain the most accurate rating system available.
Appearance-related search behaviour has moved toward AI tools across the board, including in markets with established cosmetic industries such as South Korea, and toward general-purpose assistants rather than specialist software, a pattern also visible in AI strategy adoption more widely.
About looksmaxxing.guide
looksmaxxing.guide is an independent research desk for men’s self-improvement, covering looksmaxxing, fitness, style, dating, and mindset. It sells no procedures, supplements, or coaching, takes no affiliate revenue from clinics or product brands, and labels every claim as evidence-based, clinical practice, or community-reported.
Media Contact
comms@flywheel.bz
Media Contact
Company Name: We The Flywheel
Contact Person: Thomas Prommer
Email: Send Email
Country: United States
Website: https://prommer.net
Press Release Distributed by ABNewswire.com
To view the original version on ABNewswire visit: Rate My Face ChatGPT: AI Face Scores Run 1.7 Points Too Generous