Submitted by taesiri 19 Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Cerebras 2