Scary that hallucinative behaviour also includes images, it like a new level of "lying".
I wonder how much of media like handwritten text and stuff like old vcr recordings have been transformed into AI fodder, here to hoping perhaps those areas are a new goldmine for AI training data so we get them to be even more smarter/augmentative down the line.
“once AI exhausts verifiable information, it starts trying to please people by telling them what it thinks they want to hear.”
How did this happen? Isn’t one reason to use AI to leave the emotion out of it so that you can better trust the results?
My husband’s family claims that there is a Cherokee princess in their lineage. I tried for years to prove it through Ancestry, but could not. I asked the relatives claiming it provide me proof, but they didn’t.
Then I attended an Ancestry conference where they stated that one of the most common American family history stories is that there is a Cherokee princess in the line, but there were no Cherokee princesses. It is a myth. Interesting.
So apparently now AI will make up fake sources if I ask it leading questions?
For the record, Ancestry’s guesses on the content of handwritten documents was already wrong about 20% of the time when humans were transcribing. I never trust the digital choice and always look at the original document myself.
I don’t even trust my mother’s work. We work on separate trees on the same account. I’ve made mistakes myself when in too much of a hurry and had to undo and redo long branches before.
Yes, it takes time, but the discovery is the fun part. It’s a satisfying puzzle that never ends.
Hopefully AI will give more access to more original documents more quickly. That’s all I need.
If Ancestry/AI becomes the interpreter and doesn’t let me view original documents myself, then it’s over.
That's how AI works. It's just a guessing machine. The better the training data, the better the guess, but given no data the guess will be wild. Sometimes it's wild even with data; there's a sort of randomness built in that makes up for the lack of human insight.
LLMs are token predictors. No more, no less. And there's often a lot of horsewash involved in discussions of ancestry.
Any tool aiming to provide accurate results must put in the work around the LLMs to cut off the hallucinations before they happen.
If you're using bog-standard ChatGPT, Claude, or Gemini for detailed genealogical research... well, that work hasn't been done and you're probably going to get what you paid for (nothing). They're generic chatbots, not trained researchers.
An AI system specifically constructed and trained to perform accurate analyses of the topic will most likely smoke even the best human researchers.
Scary that hallucinative behaviour also includes images, it like a new level of "lying".
I wonder how much of media like handwritten text and stuff like old vcr recordings have been transformed into AI fodder, here to hoping perhaps those areas are a new goldmine for AI training data so we get them to be even more smarter/augmentative down the line.
User: "We wuz kangz!"
AI: "Yes—you descend from an ancient line of kings. Not metaphorically, not as flattery, but as a statement of inheritance"
“once AI exhausts verifiable information, it starts trying to please people by telling them what it thinks they want to hear.”
How did this happen? Isn’t one reason to use AI to leave the emotion out of it so that you can better trust the results?
My husband’s family claims that there is a Cherokee princess in their lineage. I tried for years to prove it through Ancestry, but could not. I asked the relatives claiming it provide me proof, but they didn’t.
Then I attended an Ancestry conference where they stated that one of the most common American family history stories is that there is a Cherokee princess in the line, but there were no Cherokee princesses. It is a myth. Interesting.
So apparently now AI will make up fake sources if I ask it leading questions?
Not helping.
For the record, Ancestry’s guesses on the content of handwritten documents was already wrong about 20% of the time when humans were transcribing. I never trust the digital choice and always look at the original document myself.
I don’t even trust my mother’s work. We work on separate trees on the same account. I’ve made mistakes myself when in too much of a hurry and had to undo and redo long branches before.
Yes, it takes time, but the discovery is the fun part. It’s a satisfying puzzle that never ends.
Hopefully AI will give more access to more original documents more quickly. That’s all I need.
If Ancestry/AI becomes the interpreter and doesn’t let me view original documents myself, then it’s over.
That's how AI works. It's just a guessing machine. The better the training data, the better the guess, but given no data the guess will be wild. Sometimes it's wild even with data; there's a sort of randomness built in that makes up for the lack of human insight.
not just training data, not just fine tuning.
the quality of the context and instructions provided to the LLM is TREMENDOUSLY important.
that's why so much effort is going into making the frontier models handle large context effectively.
Very true. I've worked with them and they're pretty good about following wrong instructions.
LLMs are token predictors. No more, no less. And there's often a lot of horsewash involved in discussions of ancestry.
Any tool aiming to provide accurate results must put in the work around the LLMs to cut off the hallucinations before they happen.
If you're using bog-standard ChatGPT, Claude, or Gemini for detailed genealogical research... well, that work hasn't been done and you're probably going to get what you paid for (nothing). They're generic chatbots, not trained researchers.
An AI system specifically constructed and trained to perform accurate analyses of the topic will most likely smoke even the best human researchers.
Right. It is the fine tuning even more than the data.
The public bots are taught to reward good sounding answers over accurate.