My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
>The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash
The "some domains" are very narrow. They likely just RL'ed popular queries.
While that's annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they're very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn't use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it's incredibly clear that it has been tuned for speed and not accuracy.
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
I agree. It's much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That's fine for things like "how do I perform CPR?" but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let's see how accurate this is.](https://www.reddit.com/r/singularity/comments/1wuj72j/gemini...)
Hallucination is accurate for what I'm seeing -- e.g. it's making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
My colleague wanted to diagnose a specific error code on his car himself, and Gemini told him that it's simple to do with an OBD2 dongle - he asked it about the details thoroughly, to confirm, and bought the dongle.
It didn't work. Gemini: "Oh yeah, that obviously cannot work, it's not possible to do it through OBD2" (paraphrasing)
It was quite funny to me, but a bit less so to my colleague.
What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.