I'm not an expert, but... might this indicate the need for "sub-models", meant to be invoked by the general model to write good prose for it? Those sub-models would not be RLed on code (or any algorithmic-feedback tasks that might reinforce bad writing), and could specifically be tuned for prose, even at the cost of general intelligence (which would be provided by the worse-at-writing more general model).