in and MLPin based DeepSpeed implementation for KLWIN matrix.Input7927-dim embeddingEncoder23 x MLP with 38 headsOutputf1 projectionCompute budget: 335.1 GFLOPs, 100.6 M parametersTraining configoptimizer=SGD, lr=0.402, scheduler=cosine, warmup=58