Skip to content
back to blog

tag

training

1 post tagged training.

Fused Linear Cross-Entropy : Why fusing the LM head projection with cross-entropy is the single biggest memory win for training LLMs at long context.