Or do you? Cause if you went back and forth with ChatGPT for an hour it definitely hallucinated and lied to you at some point. Maybe consider using something else like Brilliant.org if you want to learn a topic, yanno, so you don’t propagate whatever hallucination from ChatGPT to your kids.
These days online courses like Brilliant and such are likely largely LLM generated. Are they actually vetted by experts before publication? Who knows. My money is on no, or at least, not until someone complains.
Its easy enough to prompt ChatGPT for primary sources when doing research to validate any claims its making.
You should actually do this. Ask ChatGPT to check what it said against sources. It does correct itself. The corrections are usually not major. Just mental shortcuts.
This would be a valid point maybe 3 years ago, but most chatbots will now query and verify direct sources, especially in research mode.
This is very simple to validate and verify. You could argue it may find false primary sources.
You can condemn models for a variety of other things, but acting as if this is still reality shows a lack of understanding as to modern model capabilities
Your comment is phrased as if it somehow refutes their point but it doesn't.
> Cause if you went back and forth with ChatGPT for an hour it definitely hallucinated and lied to you at some point.
If you're asserting that this is not the case today then that's going to be require pretty extraordinary evidence. "Chatbots use Google now" is not evidence that the information they provide is in fact correct.
They don't hallucinate all the time like they used to, no, but I'd be very surprised if the majority of these sorts of conversations were free of major factual errors.
I frequently notice degradation in the model model's ability to remain coherent when it searches for information online. For example I might ask Sonnet 5 "how do I build a shed" and during its search it presumably comes across an article which talks about building a shed out of paper mache, then the model responds with something like "I caution you against your plan to build a shed out of paper mache" -- Wait, what? Who said anything about building it out of paper mache?
Firstly, I can't refute a point that's little more than an opinion - my doubt is that a modern frontier model is significantly hallucinating within a relatively short conversation that can easily be verified. There is no way to refute or accept the point without a complete conversation log. I am criticizing the hyperbolic nature of the comment.
I doubt you are getting to the context level of model degredation where it reaches context limits within a verbal hour conversation.
I've just tried to recreate your example on sonnet 5, and as someone who has done DIY projects it reads completely appropriate, but I'm happy for criticism from a shed builder. It never once tells me about paper machie or creates a silly example.
This is via a prompt requesting tools and materials, and could be further improved, unfortunately, I can't paste the markdown formatting provided.
"""
Reference size used below: 8x10 ft shed. Scale material quantities to your dimensions.
Step 1: Check Regulations & Plan
Materials: None yet — just your design/plan (graph paper or free shed-plan software)
Tools: None
Skills to find:
None required, but if your shed is large or near a boundary, a quick chat with your local planning/building department saves headaches later
What to do: Confirm permit requirements, setback distances from boundaries, and max height/size allowed without permission. Sketch your design and finalize dimensions.
Step 2: Prepare the Site
Materials:
Landscape fabric (weed barrier)
Gravel/crushed stone (for drainage base, ~4-6 in depth)
Marking spray paint or stakes + string
Tools:
Shovel & spade
Wheelbarrow
Rake
Hand tamper or plate compactor
Spirit level (4 ft) or laser level
Tape measure
Builder's square (for squaring corners)
Skills to find:
Basic site leveling — not hard, but a laser level rental helps a lot if the ground has any slope
If you have poor drainage/heavy clay soil, worth asking a landscaper for advice
What to do: Clear vegetation, mark the footprint, excavate and level, add compacted gravel base for drainage.
...
"""
I won't include the whole document, can share it further but anyone can replicate just by asking sonnet
I just don't understand the need for such hyperbole, and pretending that models are still gpt3, when you can get counter evidence in seconds.
It reminds me of the craze teachers had against trusting Wikipedia - yes, you shouldn't take all claims at face value, but arguing that nothing from Wikipedia could be useful just makes the argument silly.
Err, that wasn't meant as literal example because real conversations obviously have more than 1 turn. The only thing you could prove by getting a different result is that they don't _always_ do that, even if I shared the full conversation log.
You "don't understand the need for such hyperboles" because they're not hyperboles, I don't know how you can not pick up on these errors in your own conversations.
> and pretending that models are still gpt3
I explicitly said that the new ones are better. How's that for hyperbole?
Then this entire conversation is a pointless argument - as I agree that models aren't omniscient, godlike entities that are perfect sources of truth, and that you need to use critical thinking when using them.
I don't trust models blindly, and interrogate and verify claims that they make, but that's a basic component of being a human being.
I also never try to have massive multi step conversations to the point where I'm nearing the context limits, as if there's a subclaim i need to interrogate it's far better to clear context and just start a new chat, I have notes to join up ideas.
When learning, I'm not just doing so blindly asking a model questions, I have other material up, I can look at the answer to a example question from a textbook to verify whether I have used a model to successfully learn.
This conversation is just going to devolve further into a "well it doesn't always work" to which yes, I agree, but that doesn't mean it's not useful and doesn't help the learning process.