The model is trained to follow the default template pretty closely, breaking it usually results in much worse performance and better output variance. Certain models with synthetic data in pre-training can melt down completely. At some point into this breakage you can just take the base model and it will be better.
If you want to use a custom chat scheme, use it as an overlay, don't break the default chat/tool use/reasoning template.
If you want to use a custom chat scheme, use it as an overlay, don't break the default chat/tool use/reasoning template.