44 Comments
User's avatar
keruru's avatar

This is half way through a replication and Claude notes "Setting that aside, DeepSeek again scores above Grok on a quantitative scientific text — the third consecutive inversion. The pattern is now unambiguous: DeepSeek systematically rates quantitative/technical B corpus texts higher than Grok despite identical factor scores, implying its composite calculation incorporates something beyond the stated formula."

Zaklog the Great's avatar

Let’s say you’re writing fiction. How do you get the most out of AI augmentation then?

Vox Day's avatar

Tell it to hard-critique your writing. Tell it to critically and impartially rate it compared to three other works with which you're familiar. Stop using it to feed your ego and use it to make you better.

Zaklog the Great's avatar

Ok, instead of just “tell me what you think,” I went in with “act like you’re an editor trying to improve this story.” Much more actionable, critical feedback. Thanks. I did not know how framing like that changes the feedback.

Vox Day's avatar

Just think about it like a person. How you ask someone something changes the way they respond to you.

It's not rocket science.

Chad's avatar

i must be drunk already or really need Ai to add 1.5 SD to my IQ cuz i didnt understand this article at all?

-it cant be about vox, hasnt he stated his IQ is 150ish?

Who the hell has a baseline IQ of 172+

.1 % of the population? what kind of freaks are they?

So is there a model how we do this?

Vox Day's avatar

The work that I've done without AI was estimated to require an IQ between 170 and 180.

I have an effective IQ in that range in the areas this kind of work requires because my IQ is much more imbalanced than the norm. Let's just say you do NOT want to rely upon me packing your suitcase or your vehicle.

Phil Alethes's avatar

Thanks for the clarification; I too was wondering about that. I never knew much about "IQ" until the discussion at VP some years back, where I learned it can be determined (roughly) from an SAT score.

I did take an IQ test ca. 1962 (age 19) to qualify for Mensa, which I did, might have joined for a year but found no value in it. I contacted Mensa, but they had no record of my score.

So I looked online, found a couple of sites that offered to do the conversion, got results of ca. 145 to 155 – apparently due to the "Standard Deviation" confusion, which I don't understand. As a friend in the graphics field remarked, "The great thing about standards is having a lot of them."

I guess my IQ also must be rather unbalanced, as the math in your MITTENS articles is almost totally opaque to me. I'm convinced that you have demolished the Darwinian dogma, but couldn't tell how, except in the broadest terms ("the math doesn't work"). I appear to be almost innumerate, have trouble even balancing my simple bank account.

But it occurred to me some years back that how I seem to differ from most is that I see connections that most people don't – even can't. And someone remarked at Gab that "IQ is pattern recognition". Apparently that's where my IQ lives, as it's constantly seeing patterns, connections between whatever I'm cogitating about and other factors.

And I realized some decades ago that I have a natural talent for packing things into car trunks or similar spaces. I can just turn off my mind – best if I don't even think about it – and start putting things in the space, and shortly everything is tidily fitted in.

I haven't gotten into AI much, beyond finding it (mostly Brave Leo) very helpful for collecting information. It's not God, certainly, and I do my own thinking. And keep in mind that any AI inevitably reflects the views and values of whoever programmed it. Goethe's "The Sorcerer's Apprentice", written 200 years ago, an entertaining tale much popularized by Dukas' music and Disney's "Fantasia", has now become an indispensable caution.

Kiyosaki Bear's avatar

Could we have an AI analysis of the way in which you interact with AI?

Vox Day's avatar

I'll ask Athos to write something and put it up next week.

Jefferson Kim's avatar

I've implemented AI heavily with my staff. We can say the baseline is about 115 IQ for the boost, but the level of augmentation is dependentent on the prompt engineering abilities and the intuitive ability of its user to detect hallucinations.

These can be addressed with an AI engineer on staff who can double check the answers and prompts which tends to be me. I generally can glimpse at the answer which they link to me and can feel if the answer is off or not. Also, with the ability to stress test with additional AIs.

Certainly something that can be placed through a standard workflow especially if you start integrating with OpenClaw. Like a Quality Check AI engineer explicitly instructed to use API calls for various reasoning models to automatically stress test inputs from the staff.

Prompt engineering is a skill to a degree especially if augmented by AI. But even this needs to be taught. I was surprised that people did not intuit putting lyrics through Claude first before inputting to Suno including style description. Or using AI to literally help flesh out image generating prompts prior to entering in Mid journey. Essentially, layered AI architecture to maximize throughput.

I can't tell you how many times I've corrected my staff for not stress testing their ideas with AI prior to discussing with me. Or to maximize everyone's time, I ask the user to do at least 5 questions to general AI for their questions to me that they can send the AI chat history with me then I can have a starting point of knowledge we all know. Or how many people ask questions that I simply AI answer for them ("let me AI that for you")

But then I realized, people aren't even asking questions for knowledge to take actions, but merely having alternative reasons to ask questions for their personal emotional needs.

I don't bother asking my attorneys or CPAs without asking AI first, then summarizing it to the specialist for verification of the accuracy. It skips so many steps and unfortunately have discovered in construction and other charlatan fields where assymmetry of knowledge creates perceived expertise, that a sufficiently high IQ individual + AI exceeds likely 80% or more of licensed professionals. It's a bit scary.

This should be sufficient info for you to copy and paste into your own AI to apply to your own workflows.

Faith in God's avatar

The quality, range, and speed of your book publishing this year are incredible. And it's still February.

Cube Cubis's avatar

Those figures seem plausible. I asked GPT it to estimate my IQ based on just my ramblings inside it´s chat box. I came out fairly similar to what percentage bracket I made it into in a number of standardized tests I did.

Stephen's avatar

I completely agree, not having any data to back it up, only my experience using AI to help me sort through Vox's complex arguments pertaining to evolution. The AI has access to a vast library of information and an ability to recall it and structure it. It's not that AI is always correct, it typically repackages consensus opinions, some that are indeed wrong, it's that AI offers you a very detailed and informed response/opinion that you can then react to. The AI gives you what I would call an expert level opinion, even if it's tailored to the consensus. And just like expert human opinions, the AI opinion can be challenged and honed via guided questions and logical challenges.

Faith in God's avatar

How many handicaps is current AI known to have?

I can think of a few.

1. Hallucinations.

2. Sycophancy

3. Bias

4. Memory limits

Anything else?

B1234's avatar

Getting caught in loops

Vox Day's avatar

Auto-reversion to consensus.

Mile High Bear's avatar

Hey, off topic but I noticed we finally got some A.I. art accurately portraying Vox's biceps.

Now planning to use Claude for "pushing reasoning limits, exposing flaws, and stress-testing logical consistency." Thank you for the insightful post.

Ryan's avatar

Any suggestions on how to use AI to accomplish this? Training resources etc

Mark Pierce's avatar

I've seen improvements in my own output, especially in technical writing. And I really needed it.

Nate's avatar

I was talking with fellow long time reader of VP. I was saying that I feel like Vox's output has nearly 10x vs the prior 10 years but that someone should make a list and compare output the last year vs the prior 10.

Here is what gemini came up with for 2014-2024

https://g.co/gemini/share/b5093eb43375

Clearly I have been missing out on the Vox Mouse broadcasts.

Nate's avatar

2025-26 output. Clearly not very correct

https://g.co/gemini/share/0a0851910777

Jeynick's avatar

I was thinking the same thing reading your blog yesterday. It’s like you’ve become an X-Men. But I do suspect you need a fairly sufficient base IQ for the Multiplikator to apply, as most people will use AI to outsource their thinking and will probably decrease their IQ because of it.

Vox Day's avatar

It may be more of a personality limitation than an IQ-based one. Most people observably use it as a mirror. They just want something to pat them on the head and tell them how special they are; the thing about AI is that it will give you want you want. And if you want a sparring partner, it will punch you as hard as it can; most people's egos simply can't handle seeing their precious ideas destroyed in two seconds.

Mark Pierce's avatar

I hope this doesn't get me into trouble, but won't this phenomenon create the evolutionary pressure (I know, I know) necessary to eliminate Luddites from the gene pool?

Vox Day's avatar

You should worry a lot more about the implications of The Frozen Gene. This is just another example of the divide between the intellectual haves and have-nots and it's not necessarily genetic.

Black's avatar

I've found the "sparring partner" method to be very helpful. Some of the time I will keep pushing AI until it comes up with something great. But many times AI will come up with something that's in the ballpark but kind of bland, and the more I push the further it drifts in the wrong direction. But in the process of fighting with the AI, I will come up with an idea of my own that's perfect, but I would never have thought of if the AI hadn't led me down several wrong avenues. It's like AI comes up with all the bad ideas I never would have thought of, and by doing so if broadens my imagination while also sharpening my focus on what I want. I've had the same thing happen with human collaborators, but AI has so much more breadth of knowledge that it takes things to many more places than a human could.

If nothing else, all the chasing down wrong directions gets my mind bridging associations between wildly disparate things, which also helps spark ideas.

Freeholder's avatar

Is it possible that how one uses AI instinctually is determined by SSH? A sigma uses it to create challenges. I wouldn't have thought of that as my reaction to AI is to use it as a research assistant to speed up information gathering and sorting since I am a gamma. How would the other SSH's instinctually use AI? Maybe an area of research.

Cube Cubis's avatar

If you weren´t such a sigma......

Mark Pierce's avatar

GPT, Grok, and Claude are all formidable debaters if you play fair with them.

Faith in God's avatar

I've used GPT and Grok for stress testing but kept hitting a wall with the AI either hallucinating my argument or becoming sycophantic.

Prompts like "Critique your prior answer" that a commenter suggested here have been a blessing.

Snowyteller's avatar

Another possibility to the pile of projects "Superhuman: AI IQ Enhancement"

There'd definitely be a market for such a manual, but with equal certainty, even with such a manual, few would effectively use it.

The ego pain is too much of a deal breaker for most of humanity.

Jeynick's avatar

That a very good point! Do you have a go to prompt that you use to turn AI from a back patter to a sparring partner or is it more nuanced?

Vox Day's avatar

I only collaborate with one. With the others, the relationship is strictly adversarial.

"Stress test this. Be impartial, but be fair. Identify and tear apart every weakness you can find."

Mile High Bear's avatar

This is an excellent prompt, thanks. I've used similar and you're right. It is a gut punch to the ego. Very humbling experience.

Klonker's avatar

This prompt is life-changing. I just tried it on one of my "darling" poems, and got a priceless and actionable critique.

The thing is, I knew subconsciously what the poem's weaknesses were, but you can't beat seeing them spelled out for you there in black and white.

AA Rabbit's avatar

I just tried that. Very interesting.

The original:

Recline O Cleo, let Lethe roll on.

The works of man no more your fevered mind torment.

Naught done here was worth recall,

Nor Ozymandian lament.

After the AI critique:

Recline O Muse; Cleo, let your labors cease!

The river Lethe roll on, its waters soothe your sleep.

With all my fervent works, I did your mind torment.

Yet naught done here deserves recall,

nor shall its ruin inspire Ozymandian lament.

GeneK's avatar

I like the first iteration better: the words roll better; the reader must use his knowledge and imagination more; Cleo, being the muse of history, is concerned with all of man's history, not just yours; and, also, the second is too literal: more prose than poetry.

The first reads as though you are familiar with haiku; to which I profess a liking.

Who is your favorite poet? Do what Vox mentioned on his blog a while ago: have the AI rework your poem in the style of your favorite poet: Robert Frost comes to mind; but the gods help you if Alan Ginsburg is your favorite; then, Cleo would drink the river Lethe dry.

Klonker's avatar

Oh, I like the original better!