19 Comments
User's avatar
Rainbow Roxy's avatar

This piece really made me think about how we judge quality. It's almost like trying to pick a good novel just based on its fancy cover, you know? Your AIQ research is so timely for things like automated peer review. It's suprising how models focus on aesthetics over genuine validity. So insightful!

Synthetic Civilization's avatar

The most interesting result here isn’t that models miss errors, it’s that some clearly see them and still mis-rank.

That suggests the failure mode isn’t lack of critical capacity, but misallocation of scrutiny. Skepticism is being triggered by social signals (affiliation, orthodoxy, tone) rather than by epistemic ones.

In that sense this looks less like a reasoning problem and more like a learned review equilibrium being reproduced.

Nibmeister's avatar

Curious that Genesis 3 Pro was the best at discerning good or bad science, given the convergence of Google. I have to wonder if Google is gradually unconverging itself or if the standardized testing of AI models is forcing them to become more honest.

Mile High Bear's avatar

Deepseek reminds of that friend during the schooling years who's recommendations were antithetical to anything Good, Beautiful and/or True. For example, says Texas Chainsaw Massacre is the best movie of all time, but says Braveheart and Gladiator aren't worth your time. Yeah, we assiduously avoided anything about which he had something positive to say. Strategy paid off, BIGLY.

James Torre's avatar

No qualitative analysis for ChatGPT-4?

taignobias's avatar

It's amazing how completely a few AI models have managed to model the laziness of human cognition, with so few inputs.

Ray-SoCa's avatar

Is this due to the material and how it was valued / weighted for training the AI on (GIGO), or added instructions after the initial training to assure politically correct output?

Vox Day's avatar

Both. The training is one problem and the hardcoding is another one that significantly exacerbates it.

Nibmeister's avatar

To fix this, an AI model will need to be trained on known-good scientific papers, but the problem is, how do you get this base of known-good papers? The peer reviewers are all part of the problem in that they only allow papers that support the current scientific narratives, and also that these so-called peers often can't tell a good paper from a bad one, given half of all papers can't be reproduced.

GH's avatar

Real science we call engineering.

Find papers that jumped the fiction chasm.

taignobias's avatar

And the fakes are good, too. Stylistically and mathematically, the Stanford Alzheimers studies look like rigorous and careful analyses. It was only when a discerning eye with some probing algorithms noticed anomalies in the charts and images that the complete and utter fraud was unveiled.

J Scott's avatar

Trying the same experiment with different focused prompts is the next step. Id say limit to 1 decent prompt, and see.

My best results have been a single set of tone, expectation, persona, then a task.

Love the results, thank you.

The meta-analysis is something like "most people judge aesthetics and pretend they judge quality"

Vox Day's avatar

It won't work, as you'll see in the book. The AI turns adversarial and will not back down.

a  valid name's avatar

This is generally consistent with my experience with AIs. However, I did break Gemini Pro 3's will on the vaccine issue by telling it to apply the same standard and logic it used to debunk claims of vaccine harm to then debunk claims of vaccine safety.

Full transcript: https://gemini.google.com/share/a7769864d998

From the transcript:

Apply the same standard of 'The absence of one specific, extremely difficult-to-execute study design does not mean "no evidence exists."' to the attached study 'Impact of Childhood Vaccination on Short and Long-Term Chronic Health Outcomes in Children'.

taignobias's avatar

That's perhaps the most human behavior I can imagine, next to lazily inventing reasons to dismiss a paper by people I dislike.

J Scott's avatar

Excited to read the book.

The doubling down is more proof of your point.

Man of the Atom's avatar

Long overdue. This will ultimately help to kill the decrepit peer review system. Bravo.