This is an issue self hosting as well. There’s a lot of footguns that give you slightly bad results.
I wonder what tricks one could use to ensure the model is actually performing on par with the reference api, like matching seeds or running exact benchmarks.
I wonder what tricks one could use to ensure the model is actually performing on par with the reference api, like matching seeds or running exact benchmarks.