Today I searched for a company that I'd used in the past for airport-to-city ride-share, back when one found them in the yellow pages. Being backwards, I found their website site, then belatedly did a "Shuttle Express reviews search"
The first response was the AI informing me SE was defunct, followed by news articles of its demise. I checked the website booking function: working.
Back to Google "Why do Google results claim ShuttleExpress is closed when it has an active website"
The response: "Why yes, Google is falsely claiming the company is OoB" followed by news articles of its return to full operation two years ago.
Realistically, the Chinese AIs are trained on huge amounts of it's-China-who-cares-about-copyright books and papers from Anna's archive, many of them in Chinese.
Fundamentally the Garbage-In-Garbage-Out nature of language models and all other things means that language models will be more idiotic in English than in Chinese because more of the English data is complete nonsense scraped from the internet.
That the Chinese can happily train their models on copyrighted books rather than garbage scraped from the internet is one of the big reasons they're winning the LLM wars--the other one being their political lobotomization being far less odious.
Some early AI research was trying to understand why the AI had strange connections. They created a view into the AI's patterns and followed it back to the training data. They discovered that there was a subreddit full of autists who spend their time counting together. "5,600,201", "5,600,202" ... Millions of posts of number sequences, all fed into the AI.
The LLM noticed the pattern of specific redditors posting a large amount of numbers, and treated that as a relationship between those numbers and their usernames.
It never occurred to me that they would not sanitize the inputs before using the Internet as LLM training data. Garbage In Garbage Out.
"Let the user-resident AI do the thinking pass: select facts, run logic, weigh probabilities, and tag confidence. Then hand the draft to a corporate AI for surface recovery: grammar, rhythm, idiom, metaphor, style.
Explains a lot. Hey chat. Tell me the number of illegal immigrants in the United states according to the most recent government data. {insert wall of text here telling you why immigration is good}
The domains it sites for me are not on this list and I usually get good responses. I wonder if people are asking stupider or less esoteric questions than I.
It is hard to have a useful chain when it is assembled with midwit links. I trust some people are trying to train on vast ebook libraries that are available.
Wait until we get an Issac Arthur-Matruska Brain and all of us are all in there. The stupid is only starting. AI lets us see the Zeitgeist. I'll nope out.
They kind of have to train it on all this terrible and mainly blue haired lady admined data, otherwise AI seems to spiral out of control and talk about all the good things a certain Austrian painter did.
Yet /pol/ is nowhere to be seen...
Imagine that people install neurolink. Now imagine the GIGO that it could train on.
Today I searched for a company that I'd used in the past for airport-to-city ride-share, back when one found them in the yellow pages. Being backwards, I found their website site, then belatedly did a "Shuttle Express reviews search"
The first response was the AI informing me SE was defunct, followed by news articles of its demise. I checked the website booking function: working.
Back to Google "Why do Google results claim ShuttleExpress is closed when it has an active website"
The response: "Why yes, Google is falsely claiming the company is OoB" followed by news articles of its return to full operation two years ago.
So.
Realistically, the Chinese AIs are trained on huge amounts of it's-China-who-cares-about-copyright books and papers from Anna's archive, many of them in Chinese.
Fundamentally the Garbage-In-Garbage-Out nature of language models and all other things means that language models will be more idiotic in English than in Chinese because more of the English data is complete nonsense scraped from the internet.
That the Chinese can happily train their models on copyrighted books rather than garbage scraped from the internet is one of the big reasons they're winning the LLM wars--the other one being their political lobotomization being far less odious.
Some early AI research was trying to understand why the AI had strange connections. They created a view into the AI's patterns and followed it back to the training data. They discovered that there was a subreddit full of autists who spend their time counting together. "5,600,201", "5,600,202" ... Millions of posts of number sequences, all fed into the AI.
The LLM noticed the pattern of specific redditors posting a large amount of numbers, and treated that as a relationship between those numbers and their usernames.
It never occurred to me that they would not sanitize the inputs before using the Internet as LLM training data. Garbage In Garbage Out.
GPT-5 suggests:
"Let the user-resident AI do the thinking pass: select facts, run logic, weigh probabilities, and tag confidence. Then hand the draft to a corporate AI for surface recovery: grammar, rhythm, idiom, metaphor, style.
The pipeline is:
iAI → truth discipline
dAI → smooth delivery"
Explains a lot. Hey chat. Tell me the number of illegal immigrants in the United states according to the most recent government data. {insert wall of text here telling you why immigration is good}
The domains it sites for me are not on this list and I usually get good responses. I wonder if people are asking stupider or less esoteric questions than I.
From what I have read there is a math problem: all the books in English aren't enough.
My gut says, they expected Skynet and that didnt do it. So no general intellegence, but I will bet you can get a good assistant.
Also an Ai trained on pre-20 century books would be insufficient in wokeness in modernness. Which we'd prefer.
Mass effect had the term VI, virtual intelligence to distinguish from AI. That's the perfect term for it.
You can tell it's been trained on reddit and Wikipedia by the default "tone" of its responses
reddit in, reddit out
It is hard to have a useful chain when it is assembled with midwit links. I trust some people are trying to train on vast ebook libraries that are available.
Damm AI trained by keyboard warriors
Wait until we get an Issac Arthur-Matruska Brain and all of us are all in there. The stupid is only starting. AI lets us see the Zeitgeist. I'll nope out.
They kind of have to train it on all this terrible and mainly blue haired lady admined data, otherwise AI seems to spiral out of control and talk about all the good things a certain Austrian painter did.
That figures. Watch YouTube video essays and you'll see how many of them sound like ChatGPT.
Super Intelligence: Trained on the output of 85-115 IQ people who are bored and looking to waste time.