Skip to content

deepspeed

from Orchestra-Research/AI-research-SKILLs

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

v1.0.0MIT
142
Lines
19,436
Words
9
Code Blocks

Languages

python
08-distributed-training/deepspeed/SKILL.md