LLMs are not magic boxes, they hallucinate and deviate quite a lot when asked hard to verify questions.
The only way we got a head start of using it for coding and maths was to have some formal method of validating the output as part of the training and inference time.
The only way we got a head start of using it for coding and maths was to have some formal method of validating the output as part of the training and inference time.