this post was submitted on 23 Oct 2025

579 points (99.2% liked)

Technology

75756 readers

1916 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

L4s@hackingne.ws

579

Largest study of its kind shows AI assistants misrepresent news content 45% of the time – regardless of language or territory (www.bbc.co.uk)

submitted 1 week ago by sqgl@sh.itjust.works to c/technology@lemmy.world

34 comments fedilink hide all child comments

top 34 comments

sorted by: hot top controversial new old

[–] jordanlund@lemmy.world 45 points 1 week ago (2 children)

I wish they had broke it out by AI. The article states:

"Gemini performed worst with significant issues in 76% of responses, more than double the other assistants, largely due to its poor sourcing performance."

But I don't see that anywhere in the linked PDF of the "full results".

This sort of study should also be re-done from time to time to track AI version numbers.

[–] Rothe@piefed.social 15 points 1 week ago (2 children)

It doesn't really matter, "AI" is being asked to do a task it was never meant to do. It isn't good at it, and it will never be good at it.

[–] snooggums@piefed.world 15 points 1 week ago (2 children)

Using an LLM to return accurate information is like using a shoe to hammer a nail.

[–] athatet@lemmy.zip 6 points 1 week ago

Except that a shoe is vaguely hammer ish. More like pounding a screw in with your forehead.

[–] Rooster326@programming.dev 2 points 1 week ago (1 children)

We've all done it?

[–] snooggums@piefed.world 4 points 1 week ago

Nope, my soles are too soft.

[–] Cocodapuf@lemmy.world 0 points 1 week ago* (last edited 1 week ago) (1 children)

Wow, way to completely ignore the content of the comment you're replying to. Clearly, some are better than others... so, how do the others perform? It's worth knowing before we make assertions.

The excerpt they quoted said:

"Gemini performed worst with significant issues in 76% of responses, more than double the other assistants, largely due to its poor sourcing performance."

So that implies that "the other assistants" performed more than twice as well, so presumably that means encountering serious issues less than 38% of the time (still not great, but better). But they said "more than double the other assistants", does that mean double the rate of one of the others or double the average of the others? If it's an average it would mean that some models probably performed better, while others performed worse.

This was the point, what was reported was insufficient information.

[–] Rothe@piefed.social 0 points 6 days ago (1 children)

Yes, you are a techbro. You suck because your ideas doesn't take into consideration actual real life. Fuck you.

[–] Cocodapuf@lemmy.world 1 points 6 days ago

Wow, that's just incredibly dismissive and rude. And in response to a completely reasonable comment!

Look, forget the whole AI discussion, I don't care. Here's the thing, I really like Lemmy. I really like this community and I want to continue using it as a way to have discussions with people about interesting topics. What I don't want to see is people yelling insults and swearing at any user they disagree with.

Frankly, that behavior is unwelcome. That's reddit behavior, you can go there if that's what you want to do.

[–] nick@campfyre.nickwebster.dev 2 points 1 week ago

And also which version of the models. Gemini 2.5 Flash is a completely different experience to 2.5 Pro.

[–] SaraTonin@lemmy.world 35 points 1 week ago (2 children)

There’s a few replies talking about humans misrepresenting the news. This is true, but part of the problem here is that most people understand the concept of bias - even if only to the extent of “my people neutral, your people biased”. But this is less true for LLMs. There’s research which shows that because LLMs present information authoritatively that not only do people tend to trust them, but they’re actually less likely to check the sources that the LLM provides than they would be with other forms of being presented with information.

And it’s not just news. I’ve seen people seriously argue that fringe pseudo-science is correct because they fed a very leading prompt into a chatbot and got exactly the answer they were looking for.

[–] Best_Jeanist@discuss.online 5 points 1 week ago

I wonder if people trust ChatGPT more or less than an international celebrity who is also their best friend.

[–] Axolotl_cpp@feddit.it 4 points 1 week ago

I hear a lot of people say "let's ask chatGPT" like the AI is god and know everthing 🙏, that's a big problem to be honest

[–] Yerbouti@sh.itjust.works 19 points 1 week ago

I dont understand the use people make of AI. I know a lot of of professionnal composer who are like "That's awesome, AI does the music for me now!" and I'm like, cool, now you only have the boring part of the job to do since the fun part was made by AI. Creating the music is litteraly the only fun part, I hate everything around it.

[–] paraphrand@lemmy.world 14 points 1 week ago (2 children)

Precision, nuance, and up to the moment contextual understanding are all missing from the “intelligence.”

[–] Treczoks@lemmy.world 3 points 1 week ago (1 children)

Like the average American with an 8th grade reading comprehension.

[–] snooggums@piefed.world 3 points 1 week ago

Which is what they used for the training data.

[–] FaceDeer@fedia.io 0 points 1 week ago

So it's about on par with humans, then.

[–] MonkderVierte@lemmy.zip 13 points 1 week ago (1 children)

Parrot is wrong almost half of the time. Who knew?

[–] altphoto@lemmy.today 2 points 1 week ago

Do you realize what you just said????!!!

Wow! They have reached parrot intelligence!

Next they might teach it to butterfly! You know, like you're off the ground and going somewhere in open air, but they just keep building shit right where you're flying.... And lamps!

From there, who knows?!

[–] Kissaki@feddit.org 12 points 1 week ago

Will they change their disclaimer now, from "can be wrong" to "is often wrong"? /s

[–] danc4498@lemmy.world 12 points 1 week ago

Makes sense. I have used AI for software development tasks such as manipulating SQL queries and XML files (tedious things) and am always disappointed with how AI will misinterpret some things. But it’s obvious with those when the requests fail. But for things like “the news” where there is no QA team to point out the defect, it will be much harder to notice. And when AI starts (or continues) to use AI generated posts as sources, it will get much worse.

[–] NotMyOldRedditName@lemmy.world 10 points 1 week ago

I've had someone else's AI summarize some content I created elsewhere, and it got it incredibly wrong to the point of changing the entire meaning of my original content.

[–] oplkill@lemmy.world 9 points 1 week ago

Replace CEOs by AI

[–] moistclump@lemmy.world 9 points 1 week ago (2 children)

And then I wonder how frequently humans misinterpret the mistranslated news.

[–] snooggums@piefed.world 9 points 1 week ago

Humans do it often, but they don't have billions of dollars funding their responses.

[–] Treczoks@lemmy.world 5 points 1 week ago

Worse: One third of adult actually believe the shit the AI produces.

[–] AnUnusualRelic@lemmy.world 5 points 1 week ago (1 children)

Yet the LLM seems to be what everyone is pushing, because it will supposedly get better. Haven't we reached the limits of this model and shouldn't other types of engines be tried?

[–] floofloof@lemmy.ca 4 points 1 week ago (1 children)

shouldn’t other types of engines be tried?

Sure, but the tricky bit is to be more specific than that.

[–] AnUnusualRelic@lemmy.world 3 points 1 week ago

Well, you know...

"Waves vaguely"

[–] HugeNerd@lemmy.ca 3 points 1 week ago

wrinkle: AI used for this study

[–] Jhex@lemmy.world 1 points 1 week ago

buT AI iS hERe tO StAY

[–] sin_free_for_00_days@sopuli.xyz 1 points 1 week ago

Could be better, but still a huge step up from the hate rhetoric magats get spoon fed 24/7 from Fox and friends.

[–] sirico@feddit.uk -1 points 1 week ago

So less of a percentage than the readers and mass media