Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's not like a database you can cache though, the responses are non-deterministic, the responses you get can be different the next time you query it with the exact same prompt. That's part of the point of it being generative AI (vs a question/answer system that people imagine it is).


> That's part of the point of it being generative AI (vs a question/answer system that people imagine it is).

The point I am trying to make is that not all use cases for ChatGPT are generative. There are a lot of Q&A use cases today despite the fact that these are so far beneath its true capabilities. These items could be dealt with using more economical means.

"give me a recipe for XYZ" should not require a GPU for the first turn response, much like typing in an offensive manifesto returns a boilerplate "as an AI language model..." response.

Granted, if the user then types something like "please translate the recipe to Spanish and increase the amounts by 33%", we would have to reach for the generative model. But, how many real-world users are satisfied with some simple 1-turn response and go about their day?


Sure for Bing type interfaces it makes sense but if you consider what it is more holistically there are two questions: 1) do you want it to be non-deterministic and hope the first answer is the one you want everyone to see and 2) do you invalidate the entire cache every time there's a minor update to the model? How else do you tell which keys to remove?

You'd basically be removing the entire cache every release




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: