For readers wondering, OpenRouter isn’t capable of caching as effectively as DeepSeek is because they will, for instance, switch inference providers in the middle of a session.
How can they? If you use openrouter Deepseek fine but can’t you select the model direct from Deepseek you’re basically getting api access models that way.
Fireworks directly. At the time they were the best value of cost, speed, ZDR. They got slower on me though, but I think they are retooling, so maybe things have or will get better again. I think fireworks is primarily for when you want to do your own training on top, which I wasn't doing.