With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs?
I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.
Some offer zero data retention policies, but there can be weasel words. For example, on the individual pro plan, you can turn off the setting that lets them train models on your data, but they still have a section in their terms that allows them to evaluate your anonymized data for statistical and "research" purposes. You have to actually get a signed contract along with an enterprise plan that spells out exactly what they're going to use, and what settings enable what retention.
Seems naive to think that those providers - who have a financial interest in selling the data - would not also try to weasel out of the precise definition of ‘zero’ retention.
I have no inside information, but I always assume the tickboxes that "disable ____ data" from Google/Facebook/OpenAI just disconnects it from your own account, not hides it from the provider.
There's no independent verification of what that checkbox actually does. The company can say anything, and you are unable to verify that they actually do it.
The only verification you could do so far is GDPR-style data export, and also the adherence to GDPR regulations (and even those might get skirted if they aren't operating in europe).
Aren’t they usually phrased very specifically as “we collect this data and use it to show you relevant ads, you can opt out of us showing you relevant ads”?
What about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
I worked very briefly at Baseten, and I can say that it was a perpetual annoyance (from an engineering perspective) that customers would complain about issues with their models but we couldn't actually see the inputs/outputs. I don't know about the other providers, but at Baseten they literally weren't stored anywhere.
A provider can genuinely avoid storing inputs, as the Baseten engineer below describes. That is still different from proving what code received the prompt or protecting plaintext while it runs; I built TrustedRouter to separate ZDR, attestation, and confidential routes: https://trustedrouter.com/blog/attestation-is-all-you-need?u...
> to separate ZDR, attestation, and confidential routes
Could you please clarify what that means? Given what I've been searching for, I might in principle be part of your intended customer profile, but I can't figure out whether you are merely doing routing (alternative to OpenRouter) or also inference (alternative to the names I've mentioned above). If it's merely routing, then how do you protect me from any potential misbehavior on the part of the inference provider?
Just feedback for what you're building, so please take this in a positive spirit... I'm an AI researcher and not quite an infra guy, and I'm making recommendations on token APIs for several less knowledgeable around me (I've gotten a few people set up with Baseten recently), and I couldn't figure out whether/why I would be interested in TrustedRouter. You should communicate the story better :-)
EDIT: Here's what I now understand after some digging; please correct if wrong.
There are some M token providers (not the names I listed above?) who provide cryptographic guarantees about inference services. But somebody still needs to verify what they do on each request. For an individual running a single harness, that harness would be a logical place to perform this verification if possible. For an org with N users each running their own harness, TrustedRouter solves the N*M problem and becomes the single gateway for trusted inference -- provided one somehow trusts/verifies TrustedRouter.
AWS and Azure give you the same thing for Claude and ChatGPT, no need to be stuck with open weights. They might sometimes store some of it for other purposes (I don't know the specifics), but it is emphatically not being fed back to OpenAI or Anthropic.
I would wager that’s more acceptable if said learning is not in competition with the user. If they didn’t actually produce results but created the model only, then that could be advantageous for users too. But the moment they absorb your work to sell it, or for marketing, it’s a different moral ground.
You have some secret sauce. The model trains on it. Your competitor is solving a similar problem. The model "advantageously" helps them.
Your competitor is happy and continues to pay for the subscription. Sam and Dario just resold your code.
For what it's worth LLMs still suck at reproducing my little secret algorithm/implementation while being able to solve way harder problems. I have a good guess why that's the case.
> I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs?
The services have toggles to allow prompts to be used in the training set. There is a conspiracy theory that the toggle is a false distraction and they’re actually keeping everything, and that none of the employees involved will ever whistleblow this fact.
Outside of Internet comment sections, I think most people assume these US-based companies are doing what they say.
For enterprise use there are services like AWS Bedrock which have strict isolation guarantees. There are some people who still believe those guarantees are a lie, but once someone has reached that point I don’t think they trust anything that isn’t running entirely within their house. People in that category are a very small minority, but a very vocal minority.
The impression I have (from interacting with people IRL using OpenAI and Anthropics offerings, and how they feel about the risks involved) is just the opposite. But we probably just have different life experiences.
I can name groups of people I interact with who lean both ways.
It’s still a commonly held belief that “Facebook sells your data” and it’s cool to be cynical about everything tech in many social scenes. Conceding that a tech company might be honest about something will get you classified as a bootlicker depending on who you talk to so the only winning move is to be super cynical.
Among actual professionals I work with in tech and legal, almost nobody holds a belief that these companies are blatantly lying to their customers (and zero of their employees are whistleblowing it, while said companies also have employees trying to whistleblow AI safety on Twitter daily)
Yes, I am talking about working professionals who use LLMs. Before this thread, I would have considered it surprisingly and singularly naïve if someone told me they trusted OpenAI. I still believe the common and correct take is that these companies are largely training on customer data against their consent.
I don't think they are "blatantly" lying either, just normal bog-standard lying that we've all come to accept. It's a profitable and competitive tech company.
We have already seen this lying. The toggles are opt-out, not opt-in. When you sign up, you agree to binding arbitration, which is effective for preventing lawsuits in the US. The toggles are regularly turned back on without our consent on ChatGPT and Claude. OpenAI's "don't train on my content" setting isn't even in the ChatGPT interface.
As far as I know, they haven't suffered even a tiny controversy in public opinion over any of this at all.
There's nothing to whistleblow about when it's public knowledge.
How many of the people who checked those boxes have cryptographic proof they did it? How many of those people have opted out of the arbitration clause? How many of those people would be able to claim damages? Would the amount of people who satisfy all three questions be large enough to make it worth _not_ training on user data?
I also don’t think these companies are lying at all, but I definitely think they’re training on all your data, toggle or not.
It’s truly trivial to “anonymize” and distill your prompts and model output. They could use just about any off-the-shelf cheap model for this. In fact, their TOS explicitly allows this, even with the toggle checked.
What that probably means is that the EXACT content of your prompt is secret. But the actual ideas are not. If you discover something truly novel, then yeah they get that. They can absorb trends in consumer behavior, too.
I’m sure if someone had access to all my paraphrased prompts, which retain 0% of my exact wording, they could find out literally everything about me. It’s a bit like how collecting metadata is as good (or better!) than collecting the real data.
And we all know “anonymizing” data doesn’t really exist like we think it does. Just removing names and identifiers doesn’t make anything anonymous for motivated actors. Or… say… an LLM that is trained to recognize patterns in text. Which is, like, all of them.
> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work.
They’re not saying that some prompts are mysteriously jumping into training data.
Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers.
yes this is my understanding as well, and based on [1] seems to be the case. I don't know why everyone is just believing the un-backed accusations of people probably just didn't turn off said setting (and if they did why have they not said anything to such effect)
> I don't know why everyone is just believing the un-backed accusations
Conspiratorial thinking is very common on these topics. Even bringing up the conspiracy theory about Instagram listening to your conversations and showing you related ads will bring up a surprising amount of people defending that idea on Hacker News.
Or these customers could just use AWS Bedrock...but their current CEO is an incompetent MBA unable to publicly articulate their biggest advantage, in the context of the current AI usage my companies.
You have access to all the frontier models, but...your inputs are not shared with the model vendors...neither are used to train the next model.
Why am I even doing the Amazon board job for them!??
Bedrock is really bad. It seems like they don't host the models very well because they produce tons of bugs/errors calling the model. For example you can end up with Anthropic models not returning a stop token and you end up waiting for a timeout thinking its doing something when it isn't.
Amazon is deeply invested in Anthropic and would not defame them through marketing a service whose selling point was their startup's breach of contracts.
> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model).
2023:
"The approach also aligned with the company’s broader deployment strategy, to gradually release technologies into the world for people to get used to them. Some executives, including Altman, started to parrot the same line: OpenAI needed to get the “data flywheel” going."
Anthropic happily paid billions to settle a lawsuit for pirating books. It's a trivial cost of doing business. If you're lucky you'll get a pittance after the fact by suing them, but a contract doesn't prevent them from doing the thing you don't want them to do and that they are obviously going to do given their past behaviour.
Lucky for us Apple is already alleging something to this effect in their trade secret lawsuit, so you know they'll make sure discovery turns this up if it exists.
ZDR is based on the exact same pinky-promise as training opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data.
A lot substance is hinged on the exact definition of the word "data" or "user data". In the age of post-truth everyone is claiming that they keep no "user data". Except that after running it once through some transformer program it's no longer "user data", it's something entirely else and these corpos gave ZERO promises regarding such laundered/transformed data at all, ever.
Just a thought experiment: considering training seems to be 'fair use', I wonder if they trained a tiny model to retain key info from your prompts, would mean that this would still constitute fair use, and allow them to legally claim they don't retain your data.
The guarantee on this is a (contractual) “trust me bro”, and a right to try to sue a multi-trillion-dollar company who will absolutely drive you into the ground with legal red tape.
If you are big enough to be able to withstand that, you’re already running (or trying to run) your own/open-weight models.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.